Naveen Vivek

ProjectsAgents

MCP Triage Loop

An event-driven agent that takes a failing test on a live cluster to a verified fix, a tracked issue or a root-cause fix.

Year
2026
Area
Agents
Built with
Playwright MCP, GitHub MCP, Harness, GitHub Actions

I designed and built this loop in June 2026 and presented it to Dell engineering teams. It turns “a test failed on a live cluster” into one of three clear outcomes, with a person reviewing before anything is merged.

Two ways in

  • A build failure. A GitHub Actions run fails, or a developer passes the run ID from the CLI. The agent pulls the HTML report, screenshots and network logs, then goes straight to analysis.
  • A manual report. A developer describes symptoms on a live cluster. Context is thin, so the agent reproduces the problem first.

Three steps

  1. Replicate and capture. Playwright MCP reproduces the symptoms interactively on a warm, pre-staged cluster.
  2. Analyse the codebase. The agent reads source, page objects and specs, and checks recent commits for selector or component changes.
  3. Classify the root cause and route it to the right resolution path.

Three outcomes

  • Code, config or test fix. Apply the smallest fix, deploy to the test cluster and re-run. A pass opens a pull request for human review.
  • Infrastructure or environment. Marked as not fixable by code. The agent opens a deduplicated GitHub issue with the triage report and notifies the team.
  • Flaky or performance. Fix the race or timing cause. Never hide it with a retry, because tests change cluster state.

I originally proposed building it on LangGraph. We shipped it on Harness pipelines, markdown and scripts instead, which turned out simpler for the team to own.