Building an Issue-Resolving Agent with TrueForge
Building an Issue-Resolving Agent with TrueForge (and everything that went wrong along the way)
I spent the last day of the Agent Harness Hackathon building Issue Resolver Agent — a developer operations agent that investigates a GitHub issue, diagnoses the root cause, proposes a fix, and asks for confirmation before opening a PR. It runs on TrueForge, the open-source agent harness this hackathon was built around, with code review handled by Qodo.
This is less a polished feature writeup and more an honest account of what it actually took to get a working agent submitted before the deadline — because I think that's more useful to read than a highlight reel.
The idea
I'd already built a smaller project called GitHub Issue Solver a while back — a Streamlit app that used the GitHub REST API and an LLM to help resolve issues. It worked, but it was a script with a chat window bolted on, not an agent. When I saw the hackathon's suggested pattern — "a developer operations agent" — it was an obvious fit to rebuild that idea properly: as an agent with real tools, not a wrapper around a single prompt.
The core loop I wanted:
- Fetch a real GitHub issue
- Investigate the actual repo to understand the problem
- Diagnose the root cause, not just guess at a fix
- Propose a minimal patch
- Ask before doing anything write-side — no auto-push, no auto-merge
- Open a PR once confirmed
Tech stack
- TrueForge (standalone mode) as the agent runtime
- GitHub MCP connector for real repo/issue access — the model never touches GitHub directly, everything routes through the MCP tool
- A custom
issue-resolverskill (SKILL.md) encoding the investigation workflow and guardrails - Qodo for PR review before anything merges
- Model: switched between OpenAI, Gemini, and eventually OpenRouter depending on what actually had quota available (more on that below)
What actually went wrong
Windows, immediately. npx @truefoundry/trueforge failed outright on native Windows — the ESM loader choked on a Windows-style path, and the local sandbox provider only supports macOS/Linux anyway. I didn't have time to set up WSL cleanly, so I moved the whole thing to a GitHub Codespace instead. That turned out to be a better decision than the original plan — zero local setup, and the TrueForge UI was just a forwarded port away in the browser.
Model quota roulette. I tried Gemini first and hit a limit: 0 wall on the pro model — free tier simply didn't include it. Switched to flash, still constrained. Ended up wiring in OpenRouter as a custom OpenAI-compatible provider, which worked but then ran out of credits mid-task, right after the agent had already diagnosed the issue and proposed a fix. I finished that one fix manually and said so directly in the PR description instead of pretending the whole thing was automated end to end.
A path bug that Qodo actually caught. My first real PR looked fine on the surface — but Qodo flagged a High-severity finding: the fix had been written to skills/issue-resolver/README.md instead of the actual root README.md. The issue looked resolved. It wasn't. That's exactly the kind of mistake that's easy to miss in a rushed PR and easy for a reviewer with full repo context to catch. I closed that PR, redid the fix on a clean branch, and got a clean follow-up review before merging. That whole trail — bug caught, fixed, re-reviewed — ended up being better evidence of the review process working than a clean pass on the first try would have been.
What TrueForge actually gave me
The part worth calling out: TrueForge wasn't a thin wrapper around a model call here. The GitHub MCP connector meant the agent was working against a live repo — real issues, real files, real PRs — not a simulated context. The skill system meant the investigation logic (fetch issue → locate files → classify → diagnose → propose → confirm → PR) lived as reusable instructions the harness executed, not something re-prompted from scratch every time. And the confirm-before-write guardrail was enforced by how I wrote the skill, not just a hope that the model would behave.
What I'd do differently with more time
- Wire up a sandbox (Daytona) properly so the agent could actually execute and test its proposed fixes, not just diagnose them
- Add more skills for other workflow types (feature requests, not just bugs)
- Build out subagent delegation for larger issues that touch multiple files
Final thought
The most useful part of this build wasn't the happy path — it was watching Qodo catch a mistake I'd have merged straight through otherwise. That's the actual value of the "every PR gets reviewed" rule this hackathon enforced: it's not a formality, it genuinely caught something.
Repo: https://github.com/srinidhikuchana/Issue-Resolver-Agent
Comments
Post a Comment