And then life happened... My first omnigrex:run on Omnigrex
Omnigrex had worked on a small test project.
Give it a GitHub issue, add the omnigrex:run label, and let a Developer agent prepare a pull request and a Reviewer agent review it.
Then I make the final decision.
So it was time to let Omnigrex work on itself. My hope was that, once I got through this first Work Item, I could use Omnigrex for the rest of its development. I would describe what I wanted, review the result, and let the agents handle the work in between.
I chose what looked like a small change: add a signature to comments published by agents, so I could tell which Agent Profile was speaking. There were details to get right, but the visible result was just a footer on a GitHub comment. It seemed like a reasonable first Work Item.
I added the label.
Getting to the actual work
The issue soon started collecting comments beginning with “Human handoff required.” That human was me.

This is what my first attempt at developing Omnigrex with Omnigrex looked like. “Human handoff required.” Again. 😅
The surprising part wasn’t just how many things went wrong, but how many problems I was dealing with at the same time. I was trying to get the environment working, understand why agents kept stopping, and review the code they did manage to produce. It was quite a mess.
Some failures happened before the agent could even start working. Installing the development tools needed more memory than I had allowed. After increasing it, compilation ran into a separate temporary-storage limit. Failed attempts left behind read-only cache directories that Omnigrex couldn’t clean up, which then blocked the next attempt.
Even trying again exposed a bug: after one kind of failure, re-adding the label couldn’t reactivate the agent properly. The retry needed fixing too.
Looking back, I can describe the problems separately. Working through them wasn’t nearly as tidy. Some fixes exposed another problem, but others only removed one obstacle while the rest were still there. A failed attempt could leave something behind that broke the next one, so even trying again wasn’t always a clean start.
Re-adding the label often got things moving, but I didn’t want to keep retrying until I got lucky. This was an opportunity to fix the problems while I could still see them happening.
In practice, I was running an agent locally to prepare fixes, doing manual work on the deployment host, and deploying new versions before trying again.
It finished. Except it didn’t know that.
One failure was particularly confusing. The Developer had published its changes and requested review. Then the connection to the agent ended unexpectedly.
Omnigrex treated that as a failed attempt and retried the Developer instead of moving on to the Reviewer. The agent came back, saw that the work was already done, and finished without making another change. That didn’t satisfy what Omnigrex expected either.
So I had a pull request waiting for review and a system telling me it couldn’t complete the work.
Fixing that meant making Omnigrex check what had actually completed before letting a lost connection decide the outcome. It was exactly the sort of recovery behavior I wanted the system to handle, and now I had a real example of where it didn’t.
An approval wasn’t enough
Getting the workflow moving wasn’t the only problem. I also wasn’t happy with the quality of the changes.
Omnigrex’s Reviewer approved the pull request. I ran a separate review with a stronger model on my own machine, and it found several issues. Some were about recovering correctly after interruptions; one later revision even introduced a regression that could stop Omnigrex from publishing changes.
I put the local review findings into the pull request, let Omnigrex pick them up, and reviewed the next revision. We went around that loop several times. The built-in Reviewer did request changes after seeing the findings, but it hadn’t caught them before approving.
I was starting to wonder whether I’d chosen a good enough model. Then, while writing these notes, another thought occurred to me: had I even configured its reasoning effort?
I checked. I hadn’t. Both agents were using the default, which in this setup meant minimal effort. Well, that might explain some of it.
I changed both profiles to the highest available effort for the next Work Items. Hopefully that would help, but I still had to see how much difference it would make.
Merged, with help
The pull request finally merged the next morning. Omnigrex had contributed to itself, but I had been much more involved than expected.
Much of that involvement was fixing the machinery that was supposed to make autonomous development possible: preparing environments, setting resource limits, recovering interrupted work, and deploying fixes.
That was frustrating, but it’s also part of what I want to learn from this experiment. How do you give agents enough freedom and resources to do useful work while keeping them isolated and constrained? How should the workflow behave when an agent stops halfway through, or finishes its work but loses the connection? Those questions became very concrete very quickly.
The goal is still to let Omnigrex handle application development. This first exercise showed me how much supporting engineering that goal depends on. I don’t yet know how far I’ll get, but understanding where that machinery breaks is useful progress too.
I expected agent-driven development to put more weight on the specification and the review. What I hadn’t appreciated was how much attention the operational side would need too. In this case, that meant understanding what was happening inside Omnigrex and turning failures into issues with enough evidence to fix them. I suspect that, as agents take on more implementation work, software development will need a much more automated path from production findings back to actionable issues. Whether that belongs in Omnigrex is a separate question.
Memory problems were also causing agent processes to be killed, and I still didn’t understand why. That investigation deserves its own post.
I hoped the fixes would make the next Work Item smoother, but I wasn’t ready to assume that yet. For now, I had a merged contribution and a reason to try the next one.

And eventually, this. Omnigrex’s first contribution to itself, merged. With a little more help than planned.