Get in Touch

Course Outline

Foundations of Agentic Systems in Production

  • Agentic architectures: exploring loops, tools, memory, and orchestration layers
  • The agent lifecycle: from development and deployment to continuous operation
  • Addressing the challenges of production-scale agent management

Infrastructure and Deployment Models

  • Deploying agents within containerized and cloud environments
  • Scaling patterns: comparing horizontal vs vertical scaling, along with concurrency and throttling
  • Orchestrating multiple agents and balancing workloads

Monitoring and Observability

  • Key metrics: tracking latency, success rates, memory usage, and agent call depth
  • Tracing agent activities and analyzing call graphs
  • Instrumenting observability using Prometheus, OpenTelemetry, and Grafana

Logging, Auditing, and Compliance

  • Implementing centralized logging and structured event collection
  • Ensuring compliance and auditability in agentic workflows
  • Designing audit trails and replay mechanisms for effective debugging

Performance Tuning and Resource Optimization

  • Minimizing inference overhead and optimizing agent orchestration cycles
  • Utilizing model caching and lightweight embeddings to accelerate retrieval
  • Conducting load testing and stress scenarios for AI pipelines

Cost Control and Governance

  • Analyzing agent cost drivers: API calls, memory, compute, and external integrations
  • Tracking agent-level costs and implementing chargeback models
  • Using automation policies to prevent agent sprawl and idle resource consumption

CI/CD and Rollout Strategies for Agents

  • Integrating agent pipelines into CI/CD systems
  • Testing, versioning, and establishing rollback strategies for iterative agent updates
  • Executing progressive rollouts and safe deployment mechanisms

Failure Recovery and Reliability Engineering

  • Designing for fault tolerance and graceful degradation
  • Applying retry, timeout, and circuit breaker patterns for agent reliability
  • Implementing incident response and post-mortem frameworks for AI operations

Capstone Project

  • Building and deploying an agentic AI system with full monitoring and cost tracking
  • Simulating load, measuring performance, and optimizing resource usage
  • Presenting the final architecture and monitoring dashboard to peers

Summary and Next Steps

Requirements

  • A solid grasp of MLOps and production machine learning systems
  • Practical experience with containerized deployments (Docker/Kubernetes)
  • Knowledge of cloud cost optimization and observability tools

Target Audience

  • MLOps engineers
  • Site Reliability Engineers (SREs)
  • Engineering managers overseeing AI infrastructure
 21 Hours

Testimonials (3)

Related Categories