DevOps Academy Podcast

DevOps Academy Podcast

by Ivo Radulovski
Season 1
DevOps Academy Deep Dive S1E16: eBPF and the Future of Observability — Can We Monitor Systems Without Instrumenting Everything?
AI
Modern engineering teams collect more telemetry than ever, yet getting meaningful visibility into production can still require extensive application instrumentation. In this episode of DevOps Academy Deep Dive, we explore how eBPF is changing observability by providing deeper visibility into running systems without requiring developers to modify every application. We discuss where eBPF fits alongside traditional instrumentation and OpenTelemetry, how it can help uncover service dependencies and performance problems, and why automatic observability doesn't mean collecting everything. The conversation also looks at continuous profiling, Kubernetes environments, AI-assisted troubleshooting, security, privacy, and the limits of zero-code instrumentation. 🔑 Key Takeaways What eBPF is and why it matters for DevOps and SRE teams How eBPF provides visibility without modifying application code Where traditional application instrumentation is still necessary Why technical visibility and business context are different How eBPF can help uncover real service dependencies Why it can be especially valuable for legacy and third-party applications How continuous profiling adds deeper performance context Where eBPF and OpenTelemetry complement each other How AI could correlate runtime telemetry and accelerate troubleshooting Why the future of observability is likely a combination of automatic discovery and intentional instrumentation 🎧 Tune in to explore whether the next generation of observability will require developers to instrument less while still understanding production more deeply. 💬 Could your team understand production without manually instrumenting every service?
DevOps Academy Deep Dive S1E16: When Incidents Last for Days — Are DevOps Teams Prepared for the Long Outage?
AI
Most incident response plans are built around the first few hours of an outage. But what happens when the incident continues into the next shift, the next time zone, or even the next day? In this episode of DevOps Academy Deep Dive, we explore the challenges of long-running production incidents and why they test much more than technical recovery. From responder fatigue and shift handoffs to decision logs, changing hypotheses, stakeholder communication, and organizational single points of failure, we look at what it takes to keep an incident response effective when the original team can no longer stay online. 🔑 Key Takeaways Why long-running incidents require a different response strategy How responder fatigue becomes an operational risk Why incident teams need structured shift handoffs What should be captured in a live incident document and decision log How to prevent new responders from repeating previous investigations Why incident command may need to rotate How to balance mitigation with root-cause investigation Why long outages expose knowledge and ownership gaps Where AI can help preserve incident context and prepare handoffs Why incident exercises should test people and processes, not just systems How to manage temporary emergency changes after recovery Why resilient incident response should never depend on individual heroics 🎧 Tune in to explore how DevOps and SRE teams can prepare for the incidents that don't end after the first few hours. 💬 Could your incident response process survive a complete shift change without losing critical context?
DevOps Academy Deep Dive S1E15: Who Owns Production Now? DevOps, SRE, Platform Teams—or Developers?
AI
As engineering organizations become more specialized, production ownership is getting harder to define. Should developers own what they build? Should SRE teams own reliability? Should platform teams own the runtime environment? Or should DevOps remain the connective layer across everything? In this episode of DevOps Academy Deep Dive, we explore how modern engineering teams divide responsibility for production—and where unclear ownership creates delays, bottlenecks, and incident confusion. The conversation looks at shared responsibility, service ownership, on-call, platform engineering, SRE, self-service infrastructure, Infrastructure as Code, incident response, and the growing role of AI in operational workflows. 🔑 Key Takeaways Why production ownership should be shared but never ambiguous What “you build it, you run it” really means Why central DevOps teams can become operational bottlenecks How platform engineering changes the ownership model Where SRE fits into production reliability Why application teams still need operational accountability How clear incident roles reduce confusion during outages How self-service platforms support distributed ownership How AI can automate operations without taking over responsibility Why autonomy must come with guardrails and clear ownership boundaries 🎧 Tune in to explore how modern teams can distribute production responsibility without creating uncertainty over who owns what. 💬 Who owns production in your organization today?
DevOps Academy Deep Dive S1E14: FinOps Meets DevOps — Should Engineers Be Accountable for the Cloud Bill?
AI
Cloud infrastructure is easier than ever to provision, scale, and automate. But with that flexibility comes a growing challenge: who is responsible for the cost? In this episode of DevOps Academy Deep Dive, we explore where FinOps and DevOps intersect, and why cloud economics is becoming an increasingly important part of modern engineering. The conversation looks at how engineers can balance performance, reliability, and cost without turning FinOps into a simple cost-cutting exercise. We discuss unit economics, Kubernetes cost allocation, Infrastructure as Code, platform engineering, AI workloads, observability spending, and the role of automation in identifying cloud waste. 🔑 Key Takeaways Why cloud cost is becoming an engineering concern The difference between cost reduction and cost efficiency How unit economics can make cloud spending more meaningful Why unused capacity is not always waste How overprovisioning, idle resources, storage, and telemetry increase costs Why Kubernetes can make cost attribution more difficult How Infrastructure as Code can bring cost awareness into code review 🎧 Tune in to explore why the future of DevOps is not just about shipping faster and staying reliable, but also about understanding whether infrastructure spending makes sense for the value it creates. 💬 Should DevOps engineers be responsible for the cloud bill, or should cost remain a finance concern?
DevOps Academy Deep Dive S1E13: Software Supply Chain Security — Can We Still Trust Open Source?
AI
Modern software is built on open source—but how much do we really know about the code, packages, container images, and CI/CD components entering our production environments? In this episode of DevOps Academy Deep Dive, we explore the growing challenge of software supply chain security and what DevOps teams can do to move from blind trust toward verifiable trust. The conversation looks beyond vulnerability scanning to examine the entire path from source code to production—including dependencies, compromised packages, SBOMs, artifact provenance, signing, container security, CI/CD permissions, and the emerging impact of AI-generated code. 🔑 Key Takeaways Why software dependencies have become a major security challenge How compromised and malicious packages can enter the supply chain What an SBOM tells you—and what it doesn't The difference between SBOMs and software provenance Why artifact signing matters before software reaches production How CI/CD pipelines can become a critical attack surface Why third-party actions and plugins should be treated as dependencies How DevOps can move from blind trust to verifiable trust 🎧 Tune in to explore how engineering teams can continue benefiting from open source without treating every dependency as automatically trustworthy. 💬 How confident are you that your team knows exactly what's inside the software you're shipping?
DevOps Academy Deep Dive S1E12: The Future of On-Call — Can AI Finally End Alert Fatigue?
AI
It's 2 AM, your phone goes off, and another production alert demands attention. But what if AI could determine whether you actually need to wake up? In this episode of DevOps Academy Deep Dive, we explore how AI could transform one of the most stressful parts of DevOps and Site Reliability Engineering: on-call operations and alert fatigue. From correlating hundreds of alerts into a single incident to analyzing logs, identifying recent deployments, retrieving runbooks, and recommending remediation, AI is creating new possibilities for smarter incident response. But how much control should we give it? We discuss the shift from traditional threshold-based monitoring toward intelligent, context-aware operations—and why the future may not be about eliminating on-call entirely, but ensuring engineers are interrupted only when human judgment is genuinely required. 🔑 Key Takeaways Why alert fatigue remains a major DevOps and SRE challenge How AI can correlate, deduplicate, and prioritize alerts Using AI to accelerate root-cause investigation How historical incidents and runbooks can improve AI-assisted response When autonomous remediation makes sense—and when it doesn't Why SLOs and good observability become even more important with AI How AI can identify noisy and ineffective alerts The risks of giving incident agents production access How AI could reduce on-call burnout without removing human expertise Why the future could shift from alert-driven to exception-driven operations 🎧 Tune in to discover how AI could transform incident response—and whether the 2 AM production page could eventually become the exception rather than the norm. 💬 Would you trust AI to automatically resolve a production incident without waking the on-call engineer?
DevOps Academy Deep Dive S1E11: The Skills Every DevOps Engineer Will Need by 2030
AI
What will it take to stay relevant in DevOps as we approach 2030? The DevOps role is evolving rapidly. AI agents are entering operational workflows, platform engineering is reshaping developer infrastructure, security is becoming everyone's responsibility, and engineering teams are increasingly accountable for reliability, cloud costs, and developer experience. In this episode of DevOps Academy Deep Dive, we explore the technical and human skills that could define the next generation of DevOps professionals—from strong engineering fundamentals and cloud architecture to SRE, FinOps, DevSecOps, platform engineering, and AI-assisted operations. Rather than chasing every new tool, the conversation focuses on building skills and principles that can survive the next wave of technological change. 🔑 Key Takeaways Why Linux, networking, Git, and programming fundamentals still matter How DevOps is becoming increasingly software-driven Why cloud knowledge needs to go beyond certifications The growing importance of Platform Engineering and developer experience Why FinOps and cloud cost awareness are becoming engineering skills How DevSecOps and software supply chain security are changing DevOps Why SRE, observability, resilience, and incident management matter How Infrastructure as Code will evolve alongside AI Why AI operational literacy will become essential How DevOps engineers may manage and govern autonomous AI agents Why systems thinking, communication, and adaptability will differentiate great engineers 🎧 Tune in to discover how the DevOps role could evolve by 2030—and which skills are worth investing in today. 💬 Which DevOps skill do you think will matter most by 2030?
DevOps Academy Deep Dive S1E10: Agentic AI in DevOps — Are Autonomous AI Agents Ready for Production?
AI
Autonomous AI agents are moving beyond code suggestions and chat-based assistance. They can now investigate incidents, analyze pipelines, recommend fixes, create pull requests, and, in some cases, take direct action inside operational environments. But should they be trusted with production systems? In this episode of DevOps Academy Deep Dive, we explore where agentic AI can already deliver real value in DevOps, where human approval is still essential, and what organizations need before granting AI agents access to critical infrastructure. The discussion covers controlled autonomy, least-privilege access, observability, policy as code, rollback, security risks, multi-agent systems, and the difference between an impressive AI demonstration and a production-ready operational agent. Key Takeaways How agentic AI differs from traditional DevOps automation Where autonomous agents can create immediate operational value Why unrestricted production access remains too risky How least privilege and human approval reduce blast radius What agent observability should measure Why policy enforcement must remain deterministic How teams can evaluate AI agents using real operational scenarios The safest path from read-only assistance to limited autonomy Whether agentic AI will replace or reshape DevOps roles What production readiness really means for autonomous agents Whether you work in DevOps, SRE, platform engineering, cloud operations, or engineering leadership, this episode offers a practical framework for adopting AI agents without sacrificing security, reliability, or accountability.
DevOps Academy Deep Dive S1E09: SRE vs DevOps – Where Do the Roles Converge?
AI
Is Site Reliability Engineering (SRE) replacing DevOps—or are they two sides of the same coin? In this episode of DevOps Academy Deep Dive, we explore one of the most debated topics in modern software engineering. We break down the differences between DevOps and SRE, where their responsibilities overlap, and why leading engineering organizations increasingly embrace both. From reliability engineering and automation to observability, incident management, Service Level Objectives (SLOs), and Error Budgets, this conversation explains how DevOps culture and SRE practices work together to build scalable, resilient software systems. Whether you're a DevOps Engineer, Site Reliability Engineer, Platform Engineer, Software Developer, or Engineering Leader, this episode provides practical insights into how modern engineering teams balance speed, stability, and continuous innovation. 🔑 Key Takeaways The fundamental difference between DevOps and SRE Why DevOps is a culture while SRE is an engineering discipline Understanding SLOs, SLIs, and Error Budgets How automation reduces operational toil Why observability is critical for reliable systems The growing role of Platform Engineering How AI is reshaping incident management and reliability operations Which skills engineers should develop for the future 🎧 Tune in to discover why the best engineering organizations don't choose between DevOps and SRE—they combine both. 💬 Does your organization have DevOps engineers, SREs, or a blend of both? Share your experience.
DevOps Academy Deep Dive S1E08: How Elite Engineering Teams Reduce Deployment Risk While Increasing Release Frequency
AI
Can engineering teams deploy faster without increasing risk? The world's highest-performing software organizations have proven that speed and stability don't have to compete. In this episode of DevOps Academy Deep Dive, we explore the engineering practices that enable teams to release software more frequently while maintaining reliability, resilience, and customer trust. From CI/CD automation and progressive delivery to observability, deployment strategies, and engineering culture, this conversation uncovers the principles behind modern software delivery at scale. Whether you're a DevOps engineer, Platform Engineer, SRE, Engineering Manager, or software developer, you'll discover practical insights that can help your team ship with greater confidence. 🔑 Key Takeaways Why smaller, more frequent deployments reduce risk How CI/CD builds confidence through automation and consistency The role of observability in validating production releases Why canary deployments, blue-green deployments, and feature flags matter How elite engineering teams build a culture of continuous improvement The impact of AI on deployment automation and incident response Practical steps to improve deployment reliability without slowing delivery 🎧 Tune in to learn how modern engineering teams achieve both speed and stability—and why deployment excellence is more about culture and process than tools alone. 💬 How often does your team deploy to production, and what's your biggest deployment challenge?
1 of 2