Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Basics of Audio Classification
- Types of sound events: environmental, mechanical, and human-created
- Overview of practical uses: surveillance, monitoring, and automation
- Distinguishing between audio classification, detection, and segmentation
Audio Data and Feature Extraction
- Variations in audio file types and formats
- Factors to consider regarding sampling rate, windowing, and frame size
- Extraction techniques for MFCCs, chroma features, and mel-spectrograms
Data Preparation and Annotation
- Utilising UrbanSound8K, ESC-50, and custom datasets
- Labelling sound events and their temporal boundaries
- Techniques for dataset balancing and audio augmentation
Constructing Audio Classification Models
- Application of convolutional neural networks (CNNs) in audio processing
- Model inputs: raw waveforms versus extracted features
- Selection of loss functions, evaluation metrics, and managing overfitting
Event Detection and Temporal Localisation
- Detection strategies based on frames and segments
- Post-processing detections via thresholds and smoothing techniques
- Visualising predictions across audio timelines
Advanced Concepts and Real-Time Processing
- Applying transfer learning in scenarios with limited data
- Model deployment using TensorFlow Lite or ONNX
- Considerations for streaming audio processing and latency
Project Development and Application Scenarios
- Designing a comprehensive pipeline from ingestion to classification
- Creating a proof-of-concept for surveillance, quality control, or monitoring systems
- Integration of logging, alerting, and connections to dashboards or APIs
Conclusion and Future Directions
Requirements
- Solid grasp of machine learning principles and model training processes
- Proficiency in Python programming and data pre-processing
- Basic understanding of digital audio fundamentals
Target Audience
- Data scientists
- Machine learning engineers
- Researchers and developers specialising in audio signal processing