Software Factory · 100% free setup
Your own software factory. Set up once, shipping every day.
We built a factory of AI agents that turns a plain-language request into tested, reviewed and merged code. Now we set one up around your product, your rules and your release process, so a small company can ship like a large one.
Most small companies do not have a software problem. They have a queue problem. The list of things the product should do grows every week, and the people who can change the code are too expensive to hire, too busy to start, or an agency inbox away.
We had the same queue on our own products. So instead of adding people to it, we built a factory to work it. Today it is how Geeks Invention ships changes to its internal products, and it builds new features for itself too.
This post is what that factory is, how it keeps your rules, and how we set one up for you.
Hire a team
Months to recruit, a payroll that runs whether or not there is work, and product knowledge that leaves when a person does.
The catch: slow to start, expensive to keep.
Hand it to an agency
Every change becomes a brief, an estimate and a wait. Your standards are whatever the engineer on shift happens to remember.
The catch: you pay for the waiting.
Vibe-code it yourself
Fast for a prototype. But nobody reviews it, nobody tests it, and nothing stops it ignoring how your product is meant to be built.
The catch: speed with no brakes.
A software factory keeps the speed of the third option and puts back the discipline of the first.
What it is
A line of specialist agents, and a board you can watch.
A software factory is a line of AI agents, each with one job, wrapped in machinery that decides what each of them is allowed to do. You write a request in plain words. The factory classifies it, plans it, builds it in parallel, tests it, reviews it against its acceptance criteria and merges it.
Every request becomes a work item on a board, and every work item moves through the same stages. Nobody drags the cards. A card moves only when the factory has finished a stage and proved it.
A simulated board for an invented product. The real board shows the same lanes for your work, live, along with each work item's tasks, test evidence and cost.
Inside one request
From a sentence to a merge, in six stages.
Behind each card is a run. A planner breaks the request into a handful of tasks, builders work on them side by side in separate copies of your code, and a reviewer checks the result against every acceptance criterion with evidence. Whatever fails goes back as a fix task, not as a shrug.
Intake
A small, fast model sets the type (feature, bug or refactor) and priority. The planner drafts a title, description and acceptance criteria.
Plan
The strongest model splits the work into one to six cohesive tasks, each with the files it may touch. Optional timed gate: you approve, reject, or let it approve itself.
Build
Tasks run in parallel, each in its own copy of the code. Every task loops build, test and review until its own criteria are met.
Integrate
The branch is rebased on main and your unit, integration and end-to-end suites run. A failure becomes a fix task.
Review
The reviewer marks each acceptance criterion met or unmet, with evidence. Bugs and risky changes can wait at a PR gate for you.
Merge
Machine checks pass, the change merges, and an optional smoke test runs. If the smoke test fails, the merge is reverted automatically.
Agents write code. Machinery does everything else. Branching, committing, rebasing, pushing and merging are fixed engine steps with no AI in them. No agent has a tool that can push or merge, and every command an agent asks to run executes in a sealed container with no secrets and no route to your database.
Compliance by construction
Your rules are enforced, not suggested.
A prompt that says "please follow our standards" is a hope. The factory turns your standards into checks the work cannot get past. These are the ones that matter most to a company that answers to customers, auditors or both.
Your conventions, in every prompt
Your coding conventions and the lessons from past corrections go into every agent's instructions, every time. Correct the factory once and the correction sticks.
Every task stays in its lane
Each task declares the files it may change. Anything written outside them is reverted, and a large stray change stops the work item for you to look at.
Sensitive code waits for a person
Mark payments, auth or personal data as sensitive. Any change that touches them is labelled high risk and can be held until you approve it.
No merge without proof
Five checks before any merge, approved or not: tests pass, every criterion met with evidence, scope respected, branch rebased, and your CI green.
An audit trail that cannot be edited
Every stage change writes an append-only event: who, what, when and why. Nothing is deleted, and finished work is frozen. Reopening creates new work linked to the old.
Runs on infrastructure you control
The factory is a set of containers on your own machine or server. Your code and your work history stay with you; only the model calls leave.
Built around your lifecycle
Your process, written down once.
No two companies ship the same way. Some want every bug checked by a person; some want features to merge the moment review passes. Some have a payments module nobody should touch on a Friday.
During setup we write your process into the factory's settings. After that it is not a document someone has to remember. It is how the machine behaves.
- GatesWhich changes wait for you, and for how long before they approve themselves.
- Sensitive areasThe paths that always raise the risk flag.
- Your commandsInstall, unit, integration, end-to-end, post-merge and smoke: your scripts, your CI.
- Models per roleA fast, cheap model for triage. The strongest one only for planning and review.
- LimitsTime caps, retry rounds and how many tasks run at once.
product: name: acme-billing mainBranch: main ciMode: github globs: backend: ['services/api/**'] frontend: ['apps/web/**'] tests: ['tests/**'] sensitive: ['services/api/payments/**', '**/auth/**'] commands: unit: pnpm test e2e: pnpm test:e2e smoke: ./scripts/smoke.sh gates: timeoutMin: 30 # then it approves itself planGate: true prGateTypes: ['bug'] # bugs wait for a person riskGate: hold # sensitive = always ask limits: maxRounds: 3 timeCapMin: 180 agents: classifier: { model: small-fast, effort: low } planner: { model: strongest, effort: high } reviewer: { model: strongest, effort: high }
How we set it up
One-time setup. Then autopilot.
The setup is the part that needs engineers, and it is the part we do. Once it is done, the day-to-day needs someone who knows the product, not someone who knows the code.
-
1
Map your product
Your repositories, stack, branching, test commands and how a change reaches your users today.
Geeks + you -
2
Write the rulebook
Conventions, sensitive areas, gate policy, and who approves what. Your compliance needs go here.
Geeks + you -
3
Fit the line
Agents and models per role, limits, and tests added where your code has none, so the factory has a safety net to work against.
Geeks -
4
First requests together
We run your first real requests alongside you and turn every correction into a lesson the factory keeps.
Geeks + you -
∞
You write requests. It ships.
Describe the change, answer the odd question, approve what your rules say you approve. The factory does the rest, nights and weekends included.
You
Who drives it
If you can describe it, you can ship it.
Driving the factory needs basic technical sense, not a computer science degree. The person who writes requests should know what the product is for and be able to tell whether a change works.
Your part of the week happens on one screen. Plans waiting for a nod, questions the agents would rather ask than guess, and what each piece of work cost in model usage. When an agent is unsure, it stops and asks. It does not make something up and carry on.
W-52 · Customer CSV import
3 tasks: import API, upload screen, tests. Contract: POST /customers/import
The numbers
Same AI. Organised, not improvised.
This is not a comparison with developers who do not use AI. Most teams already work with Claude Code, Codex or Cursor. The comparison is a team using those tools by hand against the same tools run as a factory. The AI stays the same. What goes is the human bottleneck: planning meetings, handoffs, review queues and merge conflicts.
Cost per month
Delivery
The honest part
What a factory does not do.
- It does not decide what to build.Choosing what your product should do next is still a business decision. The factory is very good at how, and deliberately has no opinion on what.
- It is only as safe as your tests.The checks lean on your test suites. Where those are thin, setup adds characterisation tests first, so existing behaviour is pinned down before anything changes.
- Unclear requests come back as questions.A vague request gets questions, not a guess. That costs you a minute and saves a wrong feature.
- Model usage is a real cost.It is far smaller than a salary, and it is visible per work item, so you always know what a feature cost to build.
- Some work items will stop.When rounds or time run out, the work item stops and tells you why, and the factory moves on to the next one. Nothing half-finished is ever merged.
| In-house team | Agency | Your own factory | |
|---|---|---|---|
| Time to first change | After hiring and onboarding | After a brief and an estimate | The day setup finishes |
| Follows your rules | When people remember | When it is in the contract | Every time, enforced by checks |
| Audit trail | Scattered across tools | Their system, not yours | Append-only, on your machine |
| Out of hours | Overtime | Extra cost | Same as office hours |
| Knowledge when someone leaves | Leaves with them | Stays with the agency | Stays in your rulebook and lessons |
| Cost shape | Fixed payroll | Per hour or per project | Model usage per work item |
| Example monthly cost | $25k for 5 engineers | Varies by scope | $2k for 1 junior, plus the AI you already pay for |
Get your factory set up. 100% free.
We set up the factory around your product and your process at no cost to you. It runs on the AI subscription you already pay for, so the only running cost is the one you have today.
- Your code in a Git repository
- A machine or server to run the factory on
- Your existing AI subscription: Claude Code, Codex, Cursor or similar
- A person who knows what the product should do