amarti has a real, production internal product built with zero hand-written code. I’m sure that raised an eyebrow and you’re sceptical, so I’m writing a series to explain what it is, how we stopped AI breaking things and how we built trust in the process.
AI seems to make individuals think they’re developing faster, whilst quality and reviews decline. Faros 2026 found bugs per developer up 54% and median review time up 441%.
We’re trying to combat that by adopting the dark software factory pattern, so we go faster without compromising on quality. Yes, it means we’ve almost entirely got rid of code reviews. A human still writes the specs, checks preview environments, and approves new third-party dependencies and access. Oh, and any changes to our AI guardrails.
What is a dark software factory?
Put simply, it’s where code is written entirely by AI: the plan, tests, code and implementation. A human describes what they want and checks the result, and AI does the rest.
The name comes from lights-out manufacturing, where factories run unlit because no people are needed on the floor to operate the machinery. Imagine a car being assembled in a factory by all of those robots. Now imagine it with no lights on:

Photo by Simon Kadula on Unsplash.
amarti Base is our “everything” tool internally. It handles all our HR functions, timesheets, invoicing, SOWs, expenses, ISO 27001 evidence and sales pipeline. It was written entirely by AI. There are 143k lines of AI-written code and none that a human wrote.
Choosing to build
When I started at amarti, we made the conscious choice to build rather than buy tooling like this. We wanted to avoid the experience many consultants have, where the internal tools they need to do their job are scattered across disparate and poor systems. Data ends up spread all over the place, which makes quick, effective decisions much harder.
It would have been possible to select three to five SaaS tools, negotiate and pay for each one individually, integrate them all, then build a data platform to pull the data out of each and make it available to the right places. None of the things we need amarti Base to do are that complex or difficult on their own.
Centralising data and making it available to agents in trusted ways makes it much easier to analyse and find improvements quickly. It also means we can adapt our tooling to exactly meet the needs of our business and our staff.
Why this is only possible now
Building amarti Base is a strategic investment in our consultants. It would not have been economical before AI was as effective as it is now.
I wanted to use a critical production use case to push AI to its limits and see what is possible. As a small company, our advantage is our agility, and we have used it to iterate and ship quickly.
The Faros numbers point at the problem. As PRs get cheaper to raise, the bottleneck has moved from writing the code to reviewing it. Those reviews are often done by the more senior engineers in a team, which pulls your seniors further away from creating, mentoring and coaching, and into defending and gatekeeping the quality of your codebase.
So we removed that review bottleneck, and here’s what happened.
Some stats (as of 28th September 2026)
- 143k lines of source code, trusted and in production 26 working days after implementation started.
- 12 minutes median from PR open to merge, with 90% merged within a working day.
- Only about 8% of commits were a hotfix or fix (38 of 466). There were no reverts, and none of these issues broke a critical function.
| Metric | Value |
|---|---|
| Commits in the first 4 days of development | 290 |
| Database tables | 83, built from 88 migrations |
| MCP tools | 396 (172 read, 224 write) |
| Test cases across unit, integration and e2e | 1,850 |
| AI guard and hard-constraint tests | 34 |
| Merged PRs | 266, with only 9 closed unmerged |
The issues behind that 8% of fixes:
- The production deployment ran out of memory and needed a quick fix. The live service wasn’t affected, only new deployments.
- A UX snag with the email notification logo missing in Gmail, which strips SVGs.
- Human caused: the sign-in rate limit used the wrong source for the IP address after we turned on Cloudflare protection.
- The OAuth consent-deny state denied correctly but returned the wrong reason to the client, causing an error.
- A UX snag with the people dropdown being clipped in modals, which made it hard to choose the right person.
- Human caused: production showed a LOCAL banner because the env vars were set up incorrectly by hand.
How?
This post is part of a series. Part 2 covers how we did this and stopped AI breaking things, and part 3 covers how we built trust along the way.
By