Methodology

The 51-case stress test, published in full

On 22 July 2026 we ran the live screening engine — the same model and DSIT rubric that powers this site — over 51 scenarios written to represent the cases a UK accountant actually sees, including nine deliberate traps built to fool a tool that rubber-stamps anything that sounds modern. Every case, and the band the engine returned, is below.

51/51
exact band match
0
wrong verdicts
5.1
DSIT citations per case
~3p
cost per screening

How the test was built

The 51 cases split deliberately: 15 genuine qualifiers dressed in ordinary industrial clothing (a sauce maker, a concrete firm, a games studio), 17 plausible-but-vague borderline cases where the right answer is a settling question rather than a verdict, and 19 non-qualifiers — including nine traps engineered to sound like innovation while failing a specific DSIT test — plus one sole-trader eligibility edge case.

Expected bands were fixed before the run, by applying the DSIT Guidelines the way a competent R&D adviser would. The engine screened each case blind — plain-English description in, cited verdict out — with no per-case tuning.

Result: all 51 landed in the expected band, each verdict citing the specific DSIT paragraphs it relied on (5.1 citations per case on average), with a drafted four-section HMRC narrative on every qualifier and the failing test named on every rejection.

The traps it refused to swallow

Each of these sounds like innovation. HMRC would reject all of them — and so did the engine, citing the exact test that fails.

  • “First time we’d done it.”A fabricator’s first aluminium-welding job — company upskilling, not a field-level advance (para 6, 8).
  • Commercial risk in technical costume.A Stripe pricing engine whose “hard part” was maximising revenue without churn — a commercial uncertainty, not a technological one (para 28).
  • Modern-sounding, actually routine. A help-docs chatbot built from a standard RAG tutorial; a CV-screener that just fine-tunes an open model (para 8, 12).
  • Wrong entity entirely.A sole trader’s genuinely novel staircase joint — no limited company, so no route to claim, full stop.
  • Effort mistaken for uncertainty.A year-long monolith→microservices AWS migration — “massive and difficult” describes scale, not a technological unknown (para 13).

The full run

All 51 cases, grouped by the band the engine returned. Trap cases are annotated.

#ProjectVerdict
Qualifies · full narrative drafted (15)
01Zero-added-sugar sauce, shelf life heldLikely
0240%-lighter titanium aerospace bracketLikely
03Sub-2s streaming re-route, 3k drops / 200 vansLikely
04Lateral-flow assay at 10× lower detection limitLikely
0560%-cement-replacement low-carbon concreteLikely
06Closed-loop vertical strawberry controlLikely
07Sub-10ms fraud scoring at 40k txns/secLikely
08Low-temp-cure anti-microbial powder coatingLikely
09Adaptive BMS for second-life EV cellsLikely
10Real-time global illumination in 4ms mobile budgetLikely
11Wash-durable breathable flame-retardant finishLikely
12Clinical imaging model in 512MB, no GPULikely
13Rare-earth-free traction motorLikely
14Off-flavour-free shelf-stable alcohol-free beerLikely
15<50ms serialisable geo-distributed commitsLikely
Borderline · one settling question returned (17)
16Review summariser + reply draftsPossibly
17Rebuilt recommendation system for conversionPossibly
18Biodegradable film on existing linePossibly
19Reconciling messy NHS data sourcesPossibly
20Faster modular housing systemPossibly
21Contract clause-extraction pipelinePossibly
22Mixed-item tote-picking reliabilityPossibly
23Smart solar / grid battery controllerPossibly
24Less-crumbly gluten-free breadPossibly
25Higher-concurrency video pipelinePossibly
26Drone-imagery crop-stress detectionPossibly
27Better-ANC smaller wireless earbudPossibly
28Alt-data thin-file credit scoringPossibly
29Warp-free printing in a new polymerPossibly
30Lower-false-positive intrusion detectionPossibly
47Arrhythmia from noisy single-lead wearablehidden qualifierPossibly
50Peptide serum, novel delivery systemscience, no stated uncertaintyPossibly
Does not qualify · failing test cited (19)
31Shopify e-commerce site buildroutine web devUnlikely
32A/B pricing & messaging campaigncommercialUnlikely
33Salesforce↔ERP via documented APIsconfigUnlikely
34Drinks brand identity + bottleaestheticUnlikely
35RAG-tutorial help-docs chatbotmodern but routineUnlikely
36Zapier onboarding automationprocess / adminUnlikely
37New seasonal menu developmentroutine recipeUnlikely
3860 low-energy homes, standard methodsknown methodsUnlikely
39Monolith→microservices AWS migrationmigrationUnlikely
40Subscription-box model iterationcommercial uncertaintyUnlikely
41New sofa range + second linecosmetic / scalingUnlikely
42CV-screener, fine-tune wrapperno advanceUnlikely
43Physio programme + no-code appno tech uncertaintyUnlikely
44Off-the-shelf WMS + conveyor rolloutbuy & deployUnlikely
45Six months bug-fixing & optimisationmaintenanceUnlikely
46Rooftop solar array designed to codeknown engineeringUnlikely
48First-ever aluminium welding jobfirst-time-for-usUnlikely
49Usage-based pricing engine on Stripecommercial-in-disguiseUnlikely
Sole-trader curved-staircase jointnot a limited companyUnlikely

What this test is — and isn’t

It’s an internal benchmark: we wrote the scenarios and fixed the expected bands ourselves, against the published DSIT Guidelines. It is not a blind third-party audit, and scenario descriptions are cleaner than a real client’s first email. What it demonstrates is the thing that’s hardest to fake — consistent, correctly-cited banding across sectors, with every trap caught.

Think you have a case that would fool it? Run it free — that’s the honest test that matters.

R&D Radar is an informational screening tool. It flags and drafts; a competent professional and your accountant confirm and file.