gFactory
an attempt. building in public.

by Rupreet · solo build

← All entries

Journal / Week 1 / Sunday, October 4, 2026

Finding the shape of the idea

The one assumption that decides whether gFactory is a toy: the real cost of building is not hard engineering, but the setup, babysitting and re-explaining around it.

The story

Before writing any code, I wanted to test a single claim: the thing that costs me the most time when building software isn't hard engineering. It's the setup, the babysitting and the re-explaining around the engineering.

If that's wrong, gFactory is a toy. If it's right, it's the whole thesis.

Where my time actually goes

I've already built my own agentic harness, Pebble. It's tuned and tweaked until it behaves the way I want, and it writes most of the code for me. So my work on a new project splits into three phases:

Direction. I write the PRD, the tech design and the builder specs. This takes most of my effort, and I'm fine with that. Specs and direction are where my judgment matters.

Execution. Once the specs are ready, I feed them to the harness one at a time. A big project can mean 30 to 40 specs, and each takes an hour or more to run. That's 30-plus hours of runtime, and I'm the scheduler: starting each one, checking it worked, starting the next.

Verification. Reviewing, testing, and end-to-end testing, which is still mostly unsolved, my harness included.

The pain wasn't any one phase. It was phase 2, where a human who should be thinking about direction is acting as a job queue.

The moment it became undeniable was when I started designing a context-management tool and the design sprawled to 40-plus specs. The harness could do roughly 90% of the coding. I still couldn't run the project comfortably, because I was the bottleneck between the specs and the harness.

Testing the assumption outside my own head

"I'm annoyed by this" and "this is a common bottleneck" are different claims, and only one justifies months of weekend work. So I talked to a handful of people on LinkedIn and offline.

I want to be honest about the sample: it's small, and it's conversations, not a survey. But the reaction was consistent. Nobody described algorithmic difficulty as their problem. What they wanted was a software factory: something that takes the work and produces code, without them hand-feeding it. Same shape as my frustration, from people with different stacks.

That's enough to proceed. It's not enough to claim a market, and I'd rather say so than dress it up.

The reframe: from picking agents to managing context and trust

My first mental model was an orchestration problem: which agent, which model, how they hand off to each other. But I already knew how I want to split work across models:

Specs and tech design: the top tier (Fable, Opus)

Implementation, code, reviews: the mid tier (Sonnet)

Documentation: the light tier (Haiku)

The same pattern works with the GLM family and OpenAI models, and sometimes I mix and match across providers. That routing is already a solved rule for me. It's a config choice, not a hard problem.

So the vendor is nearly interchangeable. What matters is the model tier for each type of work. The hard part is everything that has to be true before an agent's output is trustworthy enough to merge:

A shared queue that projects feed work into, so I stop being the scheduler.

Packets an agent can pick up cold, so I stop re-explaining the world to it.

A review layer that decides what's good enough to ship, with testing and e2e behind it.

That's the shift: from "which agent?" to "how does context travel, and what earns trust?"

Two constraints I set early

It has to run on subscriptions, not just API keys, and do so legally, using the supported headless modes of the tools. API-only would make it expensive to run for the kind of long, many-packet projects I'm building it for.

The review gate is non-negotiable. I'm confident about this one. Without review, testing and e2e, a factory just produces unverified code faster.

Why I picked one pilot project

My instinct was to use the factory on every project in my backlog. I dropped that, because multiple repos and stacks would test whether I can track a moving target, not whether the factory works. So the pilot is gFactory itself. The first real job I give the pipeline is to build the thing that will later run it. If it can't build its own v1, it won't build anyone else's.

What I'm not sure about

This is the part I'm most uneasy about. If gFactory works, it will shape how I build everything afterward. If its output drifts away from my vision or direction, I might not notice until late, when it's expensive to undo. I need a way to catch drift early and raise flags before it compounds. I don't have that mechanism yet, and designing it is part of the work, not an afterthought.

What got done

  • Tested the core assumption with real conversations instead of trusting my own frustration.
  • Mapped where my time goes (direction, execution, verification) and pinned the waste to execution.
  • Reframed the problem from agent orchestration to context and trust.
  • Set two constraints: works on subscriptions legally, and a mandatory review gate.
  • Chose gFactory itself as the single pilot.

Learnings

  • Separate "annoying" from "common." Only the second justifies the work.
  • The vendor is interchangeable. The model tier per task isn't. Route by type of work (design, implementation, docs), not by brand.
  • When you're the scheduler, the fix isn't faster agents. It's taking yourself out of the loop and putting a trust mechanism where you used to be.
  • A pilot that builds itself puts more pressure on a design than any number of toy examples.

Where things stand

No code, no pipeline, no specs yet. The open questions I carry into next week: is a shared, packet-based, review-gated queue the right shape or just the easiest to describe on paper, and how do I detect direction drift before it gets expensive?

Next up

  • Turn this into a scope I can commit to, then write the first spec, with gFactory writing about itself.

A question for you

If you build with coding agents, what does your equivalent of "feeding specs one by one" look like? Where do you end up acting as the scheduler?

More notes every Sunday. Browse the journal →