What Actually Happens in a Four-Week AI Build
Four weeks is not a marketing number picked because it sounds fast — it's a scope constraint, applied deliberately, before a single line of the build exists. Fix the timeline first and every scoping conversation after that has to answer one real question: what is the one capability that actually fits in three build weeks and a week of testing against real users? That constraint does more to keep a project honest than any amount of planning against an open-ended one, because there's nowhere to hide an unresolved scope disagreement when the calendar won't move.
Week one: the decisions that are expensive to reverse later
The first week is where the decisions that are cheap to change now and expensive to change in week three get made and written down — the data model, which parts of the system are deterministic logic versus a model call, where the isolation boundary sits if more than one tenant will ever use this. None of that is glamorous, and none of it is visible in a demo, which is exactly why it's tempting to defer. Deferring it is how a build looks fine in week two and then spends week three unwinding a decision that should have taken an afternoon in week one.
Weeks two and three: one capability, shipped completely
The build weeks are scoped to one core AI capability, built end to end, rather than five features each half-working — a retrieval pipeline that actually cites its sources correctly beats five integrations that each technically respond. The grounding guard and the tenant-isolation policy go in during these weeks, not after, because retrofitting either one onto a system that's already "working" means rebuilding around code that was written assuming neither existed. AI-assisted development is what makes three weeks enough time for this: the architects use it to move through implementation at the pace the timeline demands, while still writing the numbered ADR for every point where the build deviates from the agreed scope and the positive-control test for every isolation policy before it ships — the leverage buys speed on the typing, not a shortcut on the record of what got decided or the proof that it works.
Week four: user acceptance testing against real users, not assumptions
The fourth week hands the system to actual users, not to the team's own sense of whether it's ready. Fixes from that week get folded in before handover, which is the point of holding a week for it rather than shipping straight from the last build day: assumptions about how a real user will actually use the thing are exactly the assumptions a team building it can't see past on its own.
What a tight deadline is allowed to cut
When something threatens to blow the four-week timeline, the answer isn't quiet unpaid overtime and it isn't a silent under-delivery discovered at handover — it's a scoping decision, recorded as a numbered ADR the moment it's made, naming exactly what's being cut and why, agreed before the deadline arrives rather than explained after it's missed. A secondary feature, a nice-to-have integration, a wider set of edge cases than the core capability strictly needs — all fair things to trim under a fixed timeline, as long as cutting them is a decision on the record, not a surprise.
What it's never allowed to cut
Some things don't flex regardless of how the week is going, because removing them changes what's being delivered from a production system into a demo wearing production's clothes. The two architects who scoped the engagement do the actual design and review themselves, start to finish — the work doesn't quietly move to whoever has spare capacity when the schedule tightens. The grounding guard and the isolation tests ship with the core capability, not after it, because a version without them isn't a smaller version of the same system, it's a different, unproven one. And the security review before handover happens regardless of how much slack the timeline has left, because the point of four weeks was never to be fast instead of careful — it was to be both, which is the only version of "fast" worth having for a system a real user is going to depend on.
Why this holds whether you're the client or the agency reselling the work
The same structure holds for a founder shipping their first AI feature and for an agency that sold an AI capability to a client and needs a senior pod to actually build it without their own name taking the risk. Either way, the deliverable is the same: a working system, the ADRs behind every deviation from what was scoped, the isolation and grounding guarantees built in from the first commit, and a named team that's reachable after launch — not a black box that either works or doesn't, with no way to ask why.
Want a clear estimate for your project?
Book a scoping call — no commitment, just a clear-eyed read on what to build and a fixed-price estimate.
Book a scoping call