TL;DR — Key Takeaways

  • MCP gives AI assistants a consistent, typed interface to platform tools, letting them diagnose Kubernetes, observability, IaC, GitOps and CI/CD issues without engineers jumping between CLIs.
  • The biggest gains come from practical workflows such as pending-pod diagnosis, drift detection, pipeline debugging, rollbacks and incident triage.
  • Guardrails matter: read-only by default, human approval for destructive actions, per-user identity, audit logging and careful treatment of tool output as untrusted input.

Most conversations about AI in platform engineering stay abstract. This one is concrete. Over the past several months, my team has connected AI assistants to our platform through the Model Context Protocol (MCP), an open protocol that lets an AI client discover and call tools exposed by lightweight servers with typed schemas. For platform engineers, the simplest mental model is this: MCP is to AI assistants what a well-designed internal API layer is to developers. You give the model one consistent contract instead of ten CLIs.

We exposed nine domains this way: Kubernetes, metrics and logs, Infrastructure as Code, GitOps, incident management, alert routing, source control and CI/CD, workflow automation, and a small custom server for our internal deploy and rollback APIs. Below are the use cases that earned their keep, grouped by where they sit in the platform.

Kubernetes

  1. “Why is my pod stuck in Pending?” Before: An engineer runs kubectl describe, reads scheduler events, checks node taints, affinity rules and PVC bindings, then cross-references node capacity. With MCP: The developer asks the question in plain language. The agent pulls the pod spec, the scheduler events and allocatable capacity across nodes, and answers: “The pod requests 8 CPUs; no node in this pool has 8 free. Lower the request or scale the node group.” Result: Median time to diagnose Kubernetes issues dropped from about 20 minutes to under 5, and many of these questions no longer reach the platform team at all.
  2. Ingress that silently doesn’t route. Ingress failures often come down to annotations written for the wrong controller or a service with no healthy endpoints. The agent inspects the Ingress, the active controller class, the backing Service and its endpoints together, then points out the specific mismatch, such as NGINX annotations on an Ingress served by a cloud load-balancer controller, and proposes the corrected manifest.
  3. Rolling out a rotated secret: Updating a Secret or ConfigMap does not restart the pods that consume it, which is a classic source of stale-config bugs. Now an engineer can say “restart everything using db-credentials.” The agent finds every Deployment and StatefulSet that references the secret, triggers a rolling restart, and watches until all replicas are healthy on the new values.

Observability

  1. Hunting down Prometheus cardinality. Everyone knows high cardinality is killing Prometheus; finding the culprit across hundreds of targets is the hard part. The agent queries Prometheus’ own statistics, ranks metrics and labels by series count (for example, a pod-hash label generating 15,000 series), and drafts relabel rules that drop the offending label while keeping what dashboards depend on. Cardinality-driven Prometheus outages, which had hit us a few times a quarter, stopped.
  2. The dashboard that shows “No data.” A blank panel is often caused by a template variable whose source metric disappeared. The agent traces the chain from panel to variable to query to metric, finds that the metric stopped being exported, and identifies which service changed. A 30-minute investigation becomes a one-minute answer.

Infrastructure as Code and GitOps

  1. IaC that passes plan the first time. Models trained on old documentation hallucinate arguments and miss provider changes. With an IaC MCP server, the agent reads live documentation for the exact provider version pinned in the repository before writing code. Generated modules stop failing on deprecated or invented arguments.
  2. Scheduled drift detection with a human merge: A nightly agent compares IaC state against live cloud resources across workspaces. When it finds drift, such as a storage lifecycle policy changed by hand in the console, it writes the reconciling change and opens a merge request. Humans review and merge. Drift-related incidents, which had been showing up once or twice a month, became rare.
  3. “Why is my app Degraded?” When a GitOps application goes Degraded or OutOfSync, the cause is usually spread across sync history, the resource tree and events. The agent pulls all three in one pass and names the root cause, for example, a certificate that was never issued because a dependency is missing, along with the first remediation step. Application teams can now fix their own deployments without paging the platform team.

CI/CD and Deployments

  1. Pipeline failure diagnosis: The agent fetches the failed job, reads the logs, diffs the merge request against the main branch and checks pipeline variables. A typical answer: “The job fails on authentication; this branch references a CI variable that exists only in the protected environment.” Pipeline debugging went from about 20 minutes to 3 to 5.
  2. Safe, fast rollbacks through our own APIs. Our highest-leverage server is also our smallest: a few dozen lines of Python wrapping internal deploy-status and rollback endpoints. Engineers ask “what is running in production for the checkout service, and roll it back one version.” The agent shows the current and previous images, requests approval, performs the rollback and verifies health. Rollbacks dropped from 5 to 10 minutes of CLI work to under one minute, and nobody has to memorize internal endpoints.

Incident Response: Tying It Together

The individual use cases compound during incidents. When an alert fires, for instance, “error rate above 5% on the user API,” the agent pulls logs and P99 latency, checks pod health, looks for recent merges and deploys, and either takes a pre-approved low-risk action or hands the on-call engineer a complete, evidence-backed summary. Around that core, we also use MCP to:

  • Silence alerts during planned work. Before a database migration, the agent creates tightly scoped, time-boxed silences and removes them when the change completes. Deployment-driven alert storms stopped paging people.
  • Automate on-call handoffs. Open incidents, related chat threads and pending follow-ups are compiled into a handoff document for the incoming engineer.
  • Draft postmortems from facts. Timelines come straight from incident data, so the draft starts with accurate timestamps instead of memory.
  • Run multi-step runbooks. Scaling our load-test cluster, which requires resizing instances, validating readiness and deploying the workload, is now one request instead of 30 to 45 minutes of manual steps.

Across these workflows, the share of operational decisions handled without a human in the loop rose from about 5% (our old scripted automation) to roughly a third. Humans still approve anything destructive.

Lessons From Building These

  • Scope tools per task. Connecting every server to every session pushed tool definitions past 50,000 tokens. A pod investigation only needs Kubernetes and observability tools.
  • Rich errors matter. A tool that returns “plan failed” gives the model nothing to reason with. Return the full error, the diff and the affected resources.
  • Tool descriptions are prompts. The model reads your docstrings. Clear descriptions and examples prevent more mistakes than clever prompting.
  • Watch latency. Diagnostic chains are sequential, so a few hundred milliseconds per call adds up. Batch where you can.

The Guardrails Behind Every Use Case

None of this is worth doing if it widens your attack surface. The rules we treat as non-negotiable:

  1. Read-only by default. Write tools are off unless explicitly enabled per server.
  2. Human approval for destructive actions. The agent shows the exact tool, parameters and reasoning; a person approves, edits or rejects.
  3. Per-user identity. Agents act with the invoking engineer’s permissions via OAuth, with attribute-based rules on top, such as deletes allowed only in dev and staging.
  4. Tool output is untrusted input. Logs, dashboards and issue comments can carry injected instructions; strict schemas and tool risk labels limit the blast radius.
  5. Audit and observe everything. Every call is logged, and every server emits OpenTelemetry traces and metrics.
  6. Treat servers as dependencies. Scan images and review changelogs before upgrading.

Where to Start

Pick one painful, high-frequency workflow, such as Pending pods or CI failures, and connect it in read-only mode. Measure the time from question (or alert) to an actionable answer before and after. Then wrap one of your own internal APIs, because that is where your platform’s unique knowledge lives.

Platform engineering has always been about reducing cognitive load for the people who use the platform. MCP simply adds a new kind of user. Give it the same paved roads, least privilege and audit trail you give your engineers, and it becomes one of the most productive consumers of your platform.

Frequently Asked Questions

What does MCP bring to platform engineering?
MCP gives AI assistants a standard way to discover and call platform tools, similar to giving developers a consistent internal API instead of making them learn multiple command-line interfaces.
Which platform workflows benefited most?
The strongest use cases included Kubernetes troubleshooting, Prometheus cardinality analysis, IaC generation, GitOps diagnosis, CI/CD failure analysis, rollback automation and incident-response workflows.
What guardrails are important when connecting AI to platform tools?
Read-only access by default, human approval for destructive operations, per-user permissions, strict schemas, audit logging and treating all tool output as potentially untrusted.

SHARE THIS STORY