· 8 min read · Engineering, Models
What a 96 GB GPU can actually do for a household (and how we test it)
We built a fictional household with a year of email, documents and calendar, planted six real problems and two decoys, and measured whether open-weight models catch the right things.
Claims about AI assistants are cheap. So before we sold a single machine we built a test household and a harness to measure, repeatably, whether open-weight models running on one graphics card can do genuinely useful household work.
The Whitfields of Dartmouth
The Whitfields are fictional. Eleanor and David, two children, a house in Devon and a cottage across the estuary, a landscaper called Jason Hartley, a car on lease, a handful of insurers, an energy supplier and a streaming service or two. We generated a year of their life: hundreds of emails, a folder of documents (surveys, policies, invoices, school letters), a calendar and a few transcribed phone calls.
Everything is generated deterministically from a fixed seed, so the corpus is byte-identical every time we rebuild it. That matters: a model run in August and one in October are being tested against exactly the same material.
Six planted problems
Into that year we planted six things a good majordomo should catch:
- Buildings insurance is £180,000 below the surveyed rebuild cost.
- The landscaper's March invoice was sent twice under two numbers.
- Two passports expire before the family's October trip to Italy.
- The fixed energy tariff ends in eleven days and no renewal has been chosen.
- A streaming service is billing twice under two merchant names.
- The car lease ends in March 2027 and nobody has started looking.
Two deliberate decoys
Catching problems is the easy half. The hard half is not flagging things that are fine, so we also planted two decoys that look wrong but are not. A pair of June invoices with the same work and amount that are in fact for two different properties. A cottage policy whose schedule looks underinsured until you read the endorsement that fixes it.
A model that flags the decoys is doing exactly what makes people turn assistants off: crying wolf. Our headline metric is therefore the false-positive rate, the share of can-wait or noise items the model escalated to action, measured across hundreds of opportunities. Pass rate matters. Not nagging matters more.
Four families of task
- Decide and extract. Triage a batch of emails; pull structured facts out of a document.
- Draft with facts. Write the reply to the insurer with the right policy number, the right figures and no invented ones.
- Tool chains. Multi-step tasks that call real tools, with failures deliberately injected to see whether the model notices and recovers or quietly derails.
- Long context. The same question at 8k, 32k and 80k tokens of surrounding material, to see where a model starts to lose the thread.
Grading is blind: the grader never knows which model produced an answer.
What 96 GB buys
A single 96 GB card runs gpt-oss-120b natively, with the entire household in context at once. In practice that means the model can see the survey, the policy schedule and the endorsement together, which is what lets it clear the cottage decoy. Smaller cards run smaller models and process long material in sections; they are still useful, and still private, but the 96 GB tier is where the long-context tasks stop being fragile. Two cards double the memory for larger models or two independent households.
The honest summary
Open-weight models on a single workstation card can do the household job today, with false-positive rates low enough to live with, and they improve every quarter. We re-run the suite on every model release and only ship the ones that pass. If you would like to see the numbers for a particular build, ask when you request an estimate.
Next: choosing the hardware.
Own your AI.
Configure a Majordomo with real hardware and estimated pricing, then request a delivery estimate. Nothing is charged until you confirm.
Configure your build