Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News

Teller's Tech - DevOps, SRE and Cloud Podcast

Details

Ship It Weekly is a short, practical recap of what actually matters in DevOps, SRE, cloud infrastructure, and platform engineering. Each episode, your host Brian Teller walks through the latest outages, releases, tools, and incident writeups, then translates them into “here’s what this means for your systems” instead of just reading headlines. Expect a couple of main stories with context, a quick hit of tools or releases worth bookmarking, and the occasional segment on on-call, burnout, or team culture. This isn’t a certification prep show or a lab walkthrough. It’s aimed at people who are already working in the space and want to stay sharp without scrolling status pages, cloud updates, and blogs all week. You’ll hear about things like cloud provider incidents, Kubernetes and platform trends, Terraform and infrastructure changes, and real postmortems that are actually worth your time. Most episodes are 15–30 minutes, so you can catch up on the way to work or between meetings. Every now and then there will be a “special” focused on a big outage or a specific theme, but the default format is simple: what happened, why it matters, and what you might want to do about it in your own environment. If you’re the person people DM when something is broken in prod, or you’re building the cloud and platform everyone else ships on top of, Ship It Weekly is meant to be in your rotation.

Recent Episodes

OCT 1, 2026
AWS Retires DevOps Guru: What the End of Support Means, Kubernetes Cross-Namespace CVE-2026-2270, Node.js Undici WebSocket DoS & Cloudflare’s New CLI for AI Agents
This week on Ship It Weekly: AWS is retiring Amazon DevOps Guru and pointing customers toward CloudWatch and the newer Amazon DevOps Agent. Kubernetes disclosed a vulnerability where StatefulSet and ControllerRevision permissions can allow cross-namespace pod creation under specific conditions. A vulnerability in Undici can let a malicious WebSocket server crash a Node.js process through compressed data. And Cloudflare launched a new CLI as AI agents grow from 25 percent to 48 percent of Wrangler usage. The bigger theme this week is how the systems around our infrastructure are changing. Managed cloud services still have lifecycles that eventually become migration work. Kubernetes authorization can depend on what controllers do with the resources users are allowed to manipulate. Applications acting as clients still process untrusted data. And infrastructure tooling is starting to treat AI agents as first-class users rather than humans who happen to automate commands. In the lightning round: another Kubernetes vulnerability affecting Windows nodes can expose NetNTLMv2 credentials through NTLM coercion. GitHub now supports custom runners for Dependabot version and security updates. And external systems like a CMDB or internal developer portal can push repository properties into GitHub while remaining the source of truth. And the human closer comes from Lorin Hochstein and SRE Weekly. Some availability risks are probably never going away. Resources are finite, networks fail, security controls can affect availability, and production systems have to change. Preventing individual failures still matters, but incident response is part of reliability engineering too. Sometimes improving reliability means getting better at handling the failures you cannot eliminate. Links Amazon DevOps Guru End of Support - https://tsn.io/GQHN8 Kubernetes CVE-2026-2270: Cross-Namespace Pod Creation - https://tsn.io/BsNs8 Undici CVE-2026-85024: WebSocket Denial of Service - https://tsn.io/LncLd Cloudflare: Introducing the cf CLI - https://tsn.io/Wk3ma Cloudflare Forge - https://tsn.io/bAPJu Lightning Round Kubernetes CVE-2026-76654: Windows NTLM Coercion - https://tsn.io/36Kc3 GitHub: Custom Runners for Dependabot - https://tsn.io/tYl9K GitHub: External Custom Properties - https://www.tellerstech.com/go/s-1fd1396d/ Human Closer Omnipresent Availability Risks in Cloud Software - https://www.tellerstech.com/go/s-076db9dd/ Our Links This Week’s On Call Brief - https://tsn.io/fKB9V Ship It Weekly - https://tsn.io/NqkdP On Call Brief - https://tsn.io/Gpz2d
16 MIN
SEP 25, 2026
AWS Puts Elastic Beanstalk on EKS, CrowdSec Supply-Chain Breach, Critical Next.js RCE, Microsoft Disrupts EvilTokens & Why Fixing the Initial Compromise Isn’t Enough
This week on Ship It Weekly: AWS introduced Elastic Beanstalk Cluster Mode, allowing multiple applications to run on shared EKS infrastructure while AWS handles much of the Kubernetes complexity. CrowdSec published how a software supply-chain compromise led to attackers copying roughly 170 private repositories using a stolen OAuth token. A critical Next.js vulnerability in ImageResponse can lead to remote code execution through attacker-controlled SVG data. And Microsoft disrupted EvilTokens, a cybercrime platform linked to more than 12,000 compromised inboxes across 10,000 organizations. The bigger theme this week is what happens after trust has been established. Elastic Beanstalk Cluster Mode puts more infrastructure behind a managed abstraction, but shared infrastructure still means understanding isolation and blast radius. CrowdSec shows how an initial compromise can become a credential problem long after the malicious code is gone. Next.js shows how something as ordinary as generating a social preview image can expose a server-side execution path. And EvilTokens shows how attackers can use valid access to move faster once inside an account. In the lightning round: F5 has a critical BIG-IP APM vulnerability under active exploitation. GitHub Enterprise Cloud can now export an inventory of credentials with enterprise access, including PATs, SSH keys, OAuth tokens, and GitHub App credentials. Zyxel patched a vulnerability affecting GS1900 switches. And Veeam Agent for Microsoft Windows has a privilege-escalation vulnerability that can lead to SYSTEM access. And the human closer comes back to CrowdSec. Removing the malicious package, patching the server, or reimaging the workstation does not necessarily end the incident. If an attacker already stole an OAuth token, cloud credential, SSH key, session, or registry credential, that access can survive long after the original compromise is gone. Containment means understanding not only how the attacker got in, but what they took with them Links AWS Elastic Beanstalk Cluster Mode https://tsn.io/1xaV7 CrowdSec Supply-Chain Attack Analysis https://tsn.io/7yq2f Next.js ImageResponse Security Advisory https://tsn.io/8JvHp Microsoft: Disrupting EvilTokens https://tsn.io/DtbC9 Microsoft: EvilTokens and Device-Code Phishing https://tsn.io/ZzwtD F5 BIG-IP APM CVE-2026-94127 https://tsn.io/sFuKW GitHub Enterprise Credential Inventory https://tsn.io/7bpMn Zyxel GS1900 Security Advisory https://www.tellerstech.com/go/s-b2595852/ Veeam Agent for Microsoft Windows Vulnerability https://www.tellerstech.com/go/s-166d3119/ This Week’s On Call Brief https://tsn.io/Nnd8g Ship It Weekly https://tsn.io/NqkdP On Call Brief https://tsn.io/Gpz2d
17 MIN
SEP 19, 2026
GitHub Actions Security, Cisco Email Gateway RCE, Helm 3 End-of-Life, Ubuntu 26.04 Runners & Why “Nothing Changed” Is Never the Whole Story
This week on Ship It Weekly: GitHub Actions workflow execution protections are now generally available, giving organizations more control over who and what can trigger individual workflows. Cisco is patching critical vulnerabilities in Secure Email Gateway, including an actively exploited issue that can lead to remote command execution as root. Helm 3 has reached its final minor release and is heading toward end-of-life in February 2027. And GitHub’s ubuntu-latest Actions runner is preparing to move from Ubuntu 24.04 to 26.04. The bigger theme this week is infrastructure that changes even when your code does not. GitHub is making CI execution permissions more explicit, Helm teams now have a defined migration deadline, and the ubuntu-latest transition is a good example of how a completely unchanged workflow can suddenly be running in a different environment. Pinning everything forever is not necessarily the answer. The important part is knowing which dependencies are allowed to move and testing those changes deliberately. In the lightning round: GitHub Actions checks, workflow runs, and statuses will begin following your configured retention period on October 1. GitHub Advanced Security can now enforce configurations from the enterprise level. GitHub added API support for tracking when self-hosted Actions runner versions lose support. And AI Scan for pull requests can now be used without requiring CodeQL default setup. And the human closer starts with a sentence almost every infrastructure engineer has heard during an incident: “But nothing changed.” Maybe nothing changed in the application, but the runner image changed, a dependency moved, a certificate expired, DNS changed, or an external service behaved differently. Latest tags, loose version constraints, external APIs, and even support windows are dependencies. The goal is not to freeze everything forever. It is to avoid accidental mutability, where something can change without the team realizing it was ever allowed to change. Links GitHub Actions Workflow Execution Protections https://tsn.io/fbqif Cisco Secure Email Gateway Security Advisory https://tsn.io/jX2wk Helm 3 End of Life https://tsn.io/Ii7jb Ubuntu 26.04 GitHub Actions Runners and ubuntu-latest Migration https://tsn.io/7IJ9k GitHub Actions Retention Changes https://tsn.io/idFxy GitHub Advanced Security Configuration Enforcement https://tsn.io/8vRMx GitHub Actions Self-Hosted Runner Lifecycle API https://tsn.io/9UhY1 GitHub Code Scanning AI Scan https://tsn.io/ULAVW This Week’s On Call Brief https://www.tellerstech.com/go/26w38/ Ship It Weekly https://tsn.io/NqkdP On Call Brief https://tsn.io/Gpz2d
15 MIN
SEP 12, 2026
Amazon Linux 2027, GitHub Actions Cache Security, Secret-Scanning Merge Blocks, N-central CVSS 10 RCE, Karmada Graduation, ShieldCrash, CodeQL ARM64 & When Observability Fails Too
This week on Ship It Weekly: Amazon Linux 2027 enters public preview with kernel 7.1+, SELinux enforcing by default, DNF5, newer language runtimes, AWS-LC, and an x86-64-v3 baseline. GitHub Actions adds explicit cache permissions to reduce cache-poisoning risk. GitHub can now block pull requests from merging when they introduce exposed secrets. And N-able N-central has a critical pre-auth RCE that Huntress says is being actively exploited in the wild. The bigger theme this week is catching problems before they turn into incidents. Amazon Linux 2027 gives teams time to test AMIs, bootstrap scripts, agents, Terraform, CloudFormation, and CI/CD before the next platform generation becomes production reality. GitHub’s new cache controls make workflow trust boundaries explicit instead of leaving them implied. And secret-scanning rulesets move credential detection directly into the merge path, where developers can actually act on it. In the lightning round: Karmada graduates from the CNCF as multi-cluster and distributed AI scheduling grow, ShieldCrash research claims another Microsoft Defender patch bypass with SYSTEM-level access, CodeQL 2.27 adds native Linux ARM64 support, and Dependabot can now read private GitHub Packages without another personal access token. And the human closer is about what happens when observability shares the same failure domain as the thing it is watching. A full disk is bad enough. It gets worse when logs stop writing, monitoring data disappears, and the tools used to diagnose the outage start failing too. The takeaway is not that every monitoring component needs total isolation. It is that you should know what can blind you, and make sure at least one useful signal survives the failures you care about most. Links Amazon Linux 2027 Public Preview https://tsn.io/NHlEa Amazon Linux 2027 Overview and Preview Details https://tsn.io/izDYx Amazon Linux 2027 Known Issues and Preview Limitations https://tsn.io/tdugd GitHub Actions Cache Permissions with cache-mode https://tsn.io/8p94n Block Pull Requests with Exposed Secrets from Merging https://tsn.io/BspA2 N-able N-central 2026.3 Hotfix 4 https://tsn.io/xredG Huntress: N-able N-central Vulnerability and Active Exploitation https://tsn.io/QjNd5 Karmada Graduates from the CNCF https://tsn.io/lcJGh Microsoft Defender ShieldCrash Zero-Day Research https://tsn.io/YfsJ6 CodeQL 2.27 Adds Linux ARM64 Support https://tsn.io/lxfnF Automatic Dependabot Access to GitHub-Hosted Registries https://tsn.io/pSilA Ship It Weekly https://www.tellerstech.com/go/siw/ On Call Brief https://www.tellerstech.com/go/ocb/
14 MIN
SEP 4, 2026
AWS GWLB TCP Reset, Azure DevOps Live Migrations to GitHub, GitHub Runner Enforcement, Docker Root Risk, Lambda IAM Updates, PostgreSQL Upgrade Traps, SonicWall Zero-Days & Better Incident Reviews
This week on Ship It Weekly: AWS Gateway Load Balancer gets TCP Reset, giving applications a faster way to recover when firewalls or other inline appliances fail instead of waiting minutes for TCP retries to time out. Microsoft puts Enterprise Live Migrations into public preview for moving Azure DevOps repositories to GitHub Enterprise Cloud with data residency while developers keep working. GitHub is beginning enforcement against outdated self-hosted Actions runners. And Omarchy fixes a Docker configuration that effectively gave normal desktop processes a path to root. The bigger theme this week is failure modes hiding inside infrastructure we already trust. A dead network path can look like a slow application. A repository migration involves far more than copying Git history. A self-hosted runner can quietly become unsupported while it continues looking healthy. And giving a developer access to the Docker socket may sound like convenience until you remember that the Docker group is effectively a root-level privilege. In the lightning round: Lambda gets full IAM resource-based policies, AWS warns that circular PostgreSQL role memberships can stall major RDS and Aurora upgrades, a researcher releases the FalconFlank CrowdStrike privilege-escalation PoC while CrowdStrike investigates, and SonicWall patches two SMA1000 zero-days after confirming active exploitation. Links AWS Gateway Load Balancer TCP Reset https://www.tellerstech.com/go/s-d7e609ab/ Azure DevOps Enterprise Live Migrations Public Preview https://www.tellerstech.com/go/s-ea05aff9/ GitHub Actions Self-Hosted Runner Minimum Version Enforcement https://www.tellerstech.com/go/s-6e8540c4/ Omarchy: Any User Process Can Escalate to Root https://www.tellerstech.com/go/s-d22971c3/ AWS Lambda Full IAM Resource-Based Policies https://www.tellerstech.com/go/s-ff2a04b5/ Fix Circular Role Dependencies Before Upgrading RDS and Aurora PostgreSQL https://www.tellerstech.com/go/s-e4578f52/ FalconFlank CrowdStrike Privilege Escalation PoC https://www.tellerstech.com/go/s-8c21b00b/ SonicWall SMA1000 Zero-Day Advisory https://www.tellerstech.com/go/s-559ffc8b/ Remote Incident Reviews: Async First, Live Later? https://www.tellerstech.com/go/s-68ca9f5e/ This Week’s On Call Brief https://tsn.io/L95NS Ship It Weekly https://www.tellerstech.com/go/siw/ On Call Brief https://www.tellerstech.com/go/ocb/
17 MIN