The FDE Standard: 9 benchmarks for Forward Deployed Engineers
Every candidate we present has been assessed against all 9. We publish the specification in full, including the pass bar, so that you can audit our judgment rather than trust it.
Version 1.0 · Published 20 September 2026 · Revisions are logged and dated.
Production code in an unfamiliar codebase
- Definition
- Ship a working, tested change inside a repository the engineer has never seen, under time pressure.
- Why it matters
- Most engineers can build from scratch. Deployment work is almost entirely work inside somebody else’s architecture.
- Measured by
- A 4 hour live exercise in an unseen repository.
Integration under real-world constraint
- Definition
- Connect systems across authentication, rate limiting, pagination, schema drift and partial data.
- Why it matters
- Integrations fail in production on the paths nobody demonstrates. Engineers without real depth produce something that works once.
- Measured by
- An integration exercise against a deliberately unreliable interface with undocumented failure modes.
Root cause analysis in unfamiliar systems
- Definition
- Diagnose a failure in a system the engineer did not build, with incomplete logs and nobody to ask.
- Why it matters
- A Forward Deployed Engineer is usually alone at the customer site when something breaks.
- Measured by
- A timed diagnostic exercise with deliberately incomplete telemetry.
Evaluation design before implementation
- Definition
- Build the measurement harness before building the feature.
- Why it matters
- The most under-tested skill in the market, and the most common reason a deployment cannot prove its own value.
- Measured by
- The candidate receives a business requirement and must produce a runnable evaluation set with a defined threshold before writing feature code.
- Metrics
- retrieval precision@k · recall@k · MRR · nDCG
generation faithfulness · answer relevance · citation coverage · hallucination rate
end-to-end correctness · latency · cost per task · refusal rate
Retrieval system diagnosis
- Definition
- Chunking strategy, embedding selection, hybrid retrieval, reranking and context budget management, with the judgment to diagnose retrieval failure rather than blame the model.
- Why it matters
- Most systems that appear to hallucinate are retrieving the wrong context. Engineers who cannot tell the difference will try to fix it with prompts for months.
- Measured by
- A corpus with known defects deliberately planted in it.
Agent reliability under repetition
- Definition
- Tool invocation accuracy, tool selection accuracy, parameter correctness, trajectory and sequencing quality, consistency across repeated runs, robustness under paraphrase and perturbation, step and token efficiency.
- Why it matters
- An agent that works in a demonstration and fails 1 time in 5 in production is worse than no agent, because a person still has to check every output.
- Measured by
- Repeated runs against perturbed inputs, evaluated on success rate, pass@k and consistency across repetitions, following the patterns used in public benchmark suites including SWE-bench, WebArena, ToolBench and TaskBench.
Deployment and security posture
- Definition
- Containerised deployment inside the customer’s own cloud environment, single sign-on, access control, data residency, audit logging and cost observability.
- Why it matters
- Technical due diligence on data privacy and access control is a weekly activity in this role, not an edge case. An engineer who escalates every security question will stall your deployment for months.
- Measured by
- A simulated security review with a hostile reviewer.
Discovery against a mismatched brief
- Definition
- Run a discovery session where the stated problem contradicts the data the customer will actually provide.
- Why it matters
- The people who run this function at the frontier laboratories say the same thing: what the customer describes during scoping frequently does not match the reality on the ground. Catching that in week 1 rather than week 6 is most of the job.
- Measured by
- A live session with a trained stakeholder whose brief contains a planted contradiction.
Scope judgment under a failing approach
- Definition
- With a fixed deadline and an approach that will not work, decide what to cut.
- Why it matters
- The distinguishing behaviour in this role is proving out the wall fast and adjusting, rather than persisting because the plan said so.
- Measured by
- A timed exercise engineered so the obvious approach fails at roughly the halfway point.
How each benchmark is measured
Four layers, because an interview is a sample and we want a record.
- L1
Live, recorded, artefact-producing exercises. Real repositories, real unreliable interfaces, real degraded corpora. Every session recorded. Every artefact retained and available to you.
- L2
Dual-assessor rubric scoring. Two assessors score independently without seeing each other’s scores. Disagreement beyond a threshold triggers a third review. Both founders personally review every candidate who reaches a client.
- L3
Artificial-intelligence-assisted session analysis. Discovery transcripts analysed for question quality, assumption testing and whether the planted contradiction was surfaced. Code analysed for defensive patterns, error handling and test coverage, not only correctness.
- L4
Telemetry across weeks, not hours. Cohort engineers work instrumented, with written consent and full visibility of their own data. Which tools they reach for, how they use artificial intelligence in their own workflow, the depth and continuity of their focused work, and how they recover from a blocked day.
A 4 hour interview tells you how someone performs for 4 hours. Eight weeks of telemetry tells you how someone works.
ASSESSMENT SCORECARD
The benchmarks never change. The tools do.
So we test every candidate in your actual stack. These are the tools we assess in today, and the benchmark each is proven against.
Model platformsDRS-04 · DRS-06
Agent frameworksDRS-06
RetrievalDRS-05
Evaluation and observabilityDRS-04
Coding toolsDRS-01 · DRS-03
InfrastructureDRS-02 · DRS-07
DataDRS-02 · DRS-05
Product names are listed for reference. No partnership with any vendor is implied.
Revisions to the standard
| Version | Date | Change |
|---|---|---|
| 1.0 | 20 September 2026 | First publication. 9 benchmarks in 3 domains, 4 layer measurement method. |
Train to it, or hire against it.
Developers: Cohort 1 starts 16 November 2026. Companies: every candidate we present has passed all 9.