How we build

A person decides. A model drafts. More models check.

ProjectMetrix is not vibe-coded. It is not a product prompted into existence and shipped on whatever a model happened to return. Every part of it, from the domain model and the workflows to the features, the roadmap, and the individual database columns and the rules that guard them, is designed deliberately by its founder, who runs PMOs for a living, and only then built. The AI assistant drafts and implements at speed. It does not get to decide what the product is.

The judgment behind those decisions is the point. What a PMO actually needs, where projects really go wrong, and which numbers a program manager will stake a committed date on come from decades of running programs and projects, not from a model's training data. The technology accelerates the building; the experience decides what is worth building and how it has to behave.

And because any single model, including the one doing the work, has blind spots and a talent for confident mistakes, two other frontier models are brought in to attack the design before it is built. One reviews as a senior engineer, one as a security engineer, each there to be contrarian: to find the failure mode, to pressure-test the decision rather than agree with it. What survives that scrutiny gets built. The result is a product that is structurally solid, consistent and reliable by construction rather than by luck.

What "not vibe-coded" means here

Designed down to the schema

The founder's attention goes past the screens to the data model itself: the individual tables and columns, the enum that constrains a status, the row-level-security rule that decides who can read a row. If the foundation is wrong, everything an agent later reasons from is wrong, so the foundation is designed, not generated.

The roadmap is owned, not autocompleted

What gets built, in what order, and what is deliberately left out are the founder's calls, weighed against whether they compound the product's real advantage. Features are chosen because they matter to a PMO, never because they were easy for a model to produce.

Decades of delivery behind the calls

What a PMO needs, where programs actually slip, and which numbers a manager will stake a committed date on come from years of running that work, not from a model's training set. The domain judgment is lived, and it is the thing the technology is in service of.

The model is fast hands, not the architect

The AI drafts, implements and moves quickly, which is exactly what it is good at. It does not set the intent, choose the trade-offs, or decide what "correct" means. A person does that, every time, and signs off on the result.

Design is discussed before it is built

Every component, not just the big ones. The assistant proposes options with their costs and trade-offs; the product owner decides; both argue the case; and the decision — and what was rejected, and why — is written down so it is not re-litigated later. Building does not start because the previous step finished. It starts because a choice was made deliberately.

Discuss the design you + the assistant Second opinion OpenAI · Gemini Decide & record a person, in writing Build it the assistant drafts Review two AI reviewers what we learn feeds the next design
You & the assistant (Claude) Outside models — check & challenge A human decision gate
The loop that produces every component. A person is always the one who decides; the assistant never ships itself.

Two more models, before a line is written

Whenever a decision is expensive to reverse, touches tenancy or security, or could fail silently, it goes to two independent models first: OpenAI as a senior engineer (correctness, edge cases, does it match the intent) and Gemini as a security engineer (injection, tenancy, secrets, blast radius). They receive a written brief and the design docs it concerns, and they answer in prose — arguing the design, naming the failure mode, and saying what they would do instead.

Crucially, their output is evidence to weigh, not instruction to follow. The product owner reads both, keeps what survives scrutiny, and records the call. Two models reasoning independently catch what one misses — and occasionally one of them lands a genuinely important insight the partnership had not seen.

Design brief + the docs it concerns OpenAI — senior engineer correctness · edge cases · intent Gemini — security engineer injection · tenancy · secrets Weigh the evidence keep what survives Decide & build a person's call
Evidence, not instruction. Two independent models argue the design in parallel; a person weighs both and makes the call.

Reassurance where it's earned

When both outside models and the partnership converge, that's real confidence — not one model agreeing with itself. The tenant-isolation model was validated this way before it was trusted to hold customer data apart.

Blind spots, named

The change-history feature was going to store readable descriptions. A reviewer pointed out that would leak confidential detail into a portfolio-wide timeline; it now stores identifiers and resolves names under each reader's own permissions.

Insights that change the design

The assistant's activity log was set to be written in the normal request; a reviewer showed a failed request would then lose exactly the records you most want — so it's written out of band, surviving the error, and can't be forged from the app.

Every commit is reviewed by two more

Consultation happens before the build; review happens after every change — automatically. A GitHub Actions workflow runs on every push to the main branch and hands the diff to the same two models in a different role: OpenAI checking correctness, edge cases and whether the code matches the documented intent; Gemini checking for injection, tenancy leaks, auth gaps and secrets. They can't block the commit — it's already in — so instead each finding is filed automatically as a GitHub issue, which the partnership then triages and fixes together. Settled decisions are recorded alongside the workflow so the reviewers critique the code, not re-argue the design.

Push to main every change OpenAI — code review correctness · edge cases · intent Gemini — security review injection · tenancy · secrets An issue per finding nothing blocks the commit Triaged together, then fixed
The safety net after the build. Two reviewers can't block a commit, but every finding they raise is tracked and worked through.

None of this replaces judgment — it structures it. A person still owns every decision and every number the product shows. The models draft, argue, and check; the discipline is in never letting one of them, or one of us, be the only set of eyes.

Want to go deeper on how it's built?

The engineering decisions underneath the product are written up separately, and the scheduling engine is open source if you would rather read code than prose.

The engineering Get in touch Back to the product