Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Performance Concepts and Metrics
- Latency, throughput, power consumption, and resource utilisation
- Distinction between system-level and model-level bottlenecks
- Profiling methodologies for inference versus training phases
Profiling on Huawei Ascend
- Leveraging CANN Profiler and MindInsight
- Kernel and operator diagnostics
- Offload patterns and memory mapping strategies
Profiling on Biren GPU
- Performance monitoring features within the Biren SDK
- Kernel fusion, memory alignment, and execution queues
- Profiling techniques aware of power and temperature variations
Profiling on Cambricon MLU
- Performance tools including BANGPy and Neuware
- Gaining kernel-level visibility and interpreting logs
- Integrating the MLU profiler with deployment frameworks
Graph and Model-Level Optimisation
- Strategies for graph pruning and quantisation
- Operator fusion and computational graph restructuring
- Standardising input sizes and tuning batch parameters
Memory and Kernel Optimisation
- Optimising memory layout and reutilisation
- Efficient buffer management across different chipsets
- Platform-specific kernel tuning techniques
Cross-Platform Best Practices
- Performance portability through abstraction strategies
- Developing shared tuning pipelines for multi-chip environments
- Case Study: Tuning an object detection model across Ascend, Biren, and MLU
Summary and Next Steps
Requirements
- Proven experience in AI model training or deployment pipelines
- Solid understanding of GPU/MLU compute principles and model optimisation techniques
- Basic familiarity with performance profiling tools and associated metrics
Target Audience
- Performance engineers
- Machine learning infrastructure teams
- AI system architects
21 Hours