ProjectsAgents
MCP Triage Loop
An event-driven agent that takes a failing test on a live cluster to a verified fix, a tracked issue or a root-cause fix.
- Year
- 2026
- Area
- Agents
- Built with
- Playwright MCP, GitHub MCP, Harness, GitHub Actions
I designed and built this loop in June 2026 and presented it to Dell engineering teams. It turns “a test failed on a live cluster” into one of three clear outcomes, with a person reviewing before anything is merged.
Two ways in
- A build failure. A GitHub Actions run fails, or a developer passes the run ID from the CLI. The agent pulls the HTML report, screenshots and network logs, then goes straight to analysis.
- A manual report. A developer describes symptoms on a live cluster. Context is thin, so the agent reproduces the problem first.
Three steps
- Replicate and capture. Playwright MCP reproduces the symptoms interactively on a warm, pre-staged cluster.
- Analyse the codebase. The agent reads source, page objects and specs, and checks recent commits for selector or component changes.
- Classify the root cause and route it to the right resolution path.
Three outcomes
- Code, config or test fix. Apply the smallest fix, deploy to the test cluster and re-run. A pass opens a pull request for human review.
- Infrastructure or environment. Marked as not fixable by code. The agent opens a deduplicated GitHub issue with the triage report and notifies the team.
- Flaky or performance. Fix the race or timing cause. Never hide it with a retry, because tests change cluster state.
I originally proposed building it on LangGraph. We shipped it on Harness pipelines, markdown and scripts instead, which turned out simpler for the team to own.