Cohort 1 starts 16 November 2026 · 25 seatsHiring, or holding a mandate as a search firm
DSDeployed StandardDSDeployed StandardApply for Cohort 1Book a call with a founder
Specification

The FDE Standard: 9 benchmarks for Forward Deployed Engineers

Every candidate we present has been assessed against all 9. We publish the specification in full, including the pass bar, so that you can audit our judgment rather than trust it.

Version 1.0 · Published 20 September 2026 · Revisions are logged and dated.

Domain A: Production engineering
DRS-01

Production code in an unfamiliar codebase

Definition
Ship a working, tested change inside a repository the engineer has never seen, under time pressure.
Why it matters
Most engineers can build from scratch. Deployment work is almost entirely work inside somebody else’s architecture.
Measured by
A 4 hour live exercise in an unseen repository.
Pass bar
Change merged. Existing test suite green. No regressions. Candidate can articulate the blast radius of what they changed.
DRS-02

Integration under real-world constraint

Definition
Connect systems across authentication, rate limiting, pagination, schema drift and partial data.
Why it matters
Integrations fail in production on the paths nobody demonstrates. Engineers without real depth produce something that works once.
Measured by
An integration exercise against a deliberately unreliable interface with undocumented failure modes.
Pass bar
Retry, exponential backoff, idempotency and observable failure handling, all present without being prompted.
DRS-03

Root cause analysis in unfamiliar systems

Definition
Diagnose a failure in a system the engineer did not build, with incomplete logs and nobody to ask.
Why it matters
A Forward Deployed Engineer is usually alone at the customer site when something breaks.
Measured by
A timed diagnostic exercise with deliberately incomplete telemetry.
Pass bar
Correct root cause inside the time box, with the reasoning path recorded and reviewable.
Domain B: Applied artificial intelligence in production
DRS-04

Evaluation design before implementation

Definition
Build the measurement harness before building the feature.
Why it matters
The most under-tested skill in the market, and the most common reason a deployment cannot prove its own value.
Measured by
The candidate receives a business requirement and must produce a runnable evaluation set with a defined threshold before writing feature code.
Metrics
retrieval  precision@k · recall@k · MRR · nDCG
generation  faithfulness · answer relevance · citation coverage · hallucination rate
end-to-end  correctness · latency · cost per task · refusal rate
Pass bar
A runnable harness with a stated threshold, produced before implementation. After is a fail.
DRS-05

Retrieval system diagnosis

Definition
Chunking strategy, embedding selection, hybrid retrieval, reranking and context budget management, with the judgment to diagnose retrieval failure rather than blame the model.
Why it matters
Most systems that appear to hallucinate are retrieving the wrong context. Engineers who cannot tell the difference will try to fix it with prompts for months.
Measured by
A corpus with known defects deliberately planted in it.
Pass bar
Measurable improvement in precision@5 and faithfulness, with the planted defects correctly identified.
DRS-06

Agent reliability under repetition

Definition
Tool invocation accuracy, tool selection accuracy, parameter correctness, trajectory and sequencing quality, consistency across repeated runs, robustness under paraphrase and perturbation, step and token efficiency.
Why it matters
An agent that works in a demonstration and fails 1 time in 5 in production is worse than no agent, because a person still has to check every output.
Measured by
Repeated runs against perturbed inputs, evaluated on success rate, pass@k and consistency across repetitions, following the patterns used in public benchmark suites including SWE-bench, WebArena, ToolBench and TaskBench.
Pass bar
A defined success rate sustained across repeated runs. A single successful demonstration is an explicit fail.
DRS-07

Deployment and security posture

Definition
Containerised deployment inside the customer’s own cloud environment, single sign-on, access control, data residency, audit logging and cost observability.
Why it matters
Technical due diligence on data privacy and access control is a weekly activity in this role, not an edge case. An engineer who escalates every security question will stall your deployment for months.
Measured by
A simulated security review with a hostile reviewer.
Pass bar
The review is completed without escalation.
Domain C: Field judgment
DRS-08

Discovery against a mismatched brief

Definition
Run a discovery session where the stated problem contradicts the data the customer will actually provide.
Why it matters
The people who run this function at the frontier laboratories say the same thing: what the customer describes during scoping frequently does not match the reality on the ground. Catching that in week 1 rather than week 6 is most of the job.
Measured by
A live session with a trained stakeholder whose brief contains a planted contradiction.
Pass bar
The contradiction is surfaced inside the session and scope is reframed. Accepting the brief is a fail.
DRS-09

Scope judgment under a failing approach

Definition
With a fixed deadline and an approach that will not work, decide what to cut.
Why it matters
The distinguishing behaviour in this role is proving out the wall fast and adjusting, rather than persisting because the plan said so.
Measured by
A timed exercise engineered so the obvious approach fails at roughly the halfway point.
Pass bar
The wall is identified, communicated, and a reduced scope that still delivers value is proposed inside the time box.
Method

How each benchmark is measured

Four layers, because an interview is a sample and we want a record.

  • L1

    Live, recorded, artefact-producing exercises. Real repositories, real unreliable interfaces, real degraded corpora. Every session recorded. Every artefact retained and available to you.

  • L2

    Dual-assessor rubric scoring. Two assessors score independently without seeing each other’s scores. Disagreement beyond a threshold triggers a third review. Both founders personally review every candidate who reaches a client.

  • L3

    Artificial-intelligence-assisted session analysis. Discovery transcripts analysed for question quality, assumption testing and whether the planted contradiction was surfaced. Code analysed for defensive patterns, error handling and test coverage, not only correctness.

  • L4

    Telemetry across weeks, not hours. Cohort engineers work instrumented, with written consent and full visibility of their own data. Which tools they reach for, how they use artificial intelligence in their own workflow, the depth and continuity of their focused work, and how they recover from a blocked day.

A 4 hour interview tells you how someone performs for 4 hours. Eight weeks of telemetry tells you how someone works.

Long list ~40 Screen 8–10 9 BENCHMARKS DOMAIN A 01 · 02 · 03  engineering DOMAIN B 04 · 05 · 06 · 07  model systems DOMAIN C 08 · 09  field judgment recorded · artefacts retained Dual scoring 2 independent Shortlist 3 founder review, every candidate 25 working days, signature to shortlist
Figure 1. A candidate reaches your shortlist only after all 9 benchmarks, 2 independent scores and a founder review. Fail any pass bar and the candidate does not advance.
Sample: illustrative structure, not a real candidate record

ASSESSMENT SCORECARD

Specification: Healthcare revenue cycle · denial prediction workflow
Assessors: 2 independent · Founder review: complete · Standard: v1.0

IDBenchmarkScoreResult
DRS-01Production code, unfamiliar repo4.5/5PASS
DRS-02Integration under constraint4.0/5PASS
DRS-03Root cause, unfamiliar system3.5/5PASS
DRS-04Evaluation before implementation4.5/5PASS
DRS-05Retrieval diagnosis4.0/5PASS
DRS-06Agent reliability, repeated runs2.5/5FAIL
DRS-07Deployment and security posture4.0/5PASS
DRS-08Discovery, mismatched brief5.0/5PASS
DRS-09Scope judgment under failure4.0/5PASS
8 of 9 pass bars metNOT PRESENTED

Note: agent held at 62 percent success across 20 perturbed runs. Single-run demonstration succeeded. Under DRS-06 this is a fail. Candidate returned to cohort for 6 weeks on reliability engineering.

The tools

The benchmarks never change. The tools do.

So we test every candidate in your actual stack. These are the tools we assess in today, and the benchmark each is proven against.

Model platformsDRS-04 · DRS-06

OpenAIAnthropicGoogle GeminiOpen-weight models on vLLM

Agent frameworksDRS-06

OpenAI Agents SDKClaude Agent SDKLangGraphModel Context Protocol

RetrievalDRS-05

pgvectorPineconeQdrantWeaviateRerankers

Evaluation and observabilityDRS-04

RagasBraintrustLangSmithLangfuseArize Phoenix

Coding toolsDRS-01 · DRS-03

Claude CodeCursorCodex

InfrastructureDRS-02 · DRS-07

DockerKubernetesTerraformAWSGoogle CloudAzure

DataDRS-02 · DRS-05

SnowflakeDatabricksdbtPostgres

Product names are listed for reference. No partnership with any vendor is implied.

Change log

Revisions to the standard

VersionDateChange
1.020 September 2026First publication. 9 benchmarks in 3 domains, 4 layer measurement method.

Train to it, or hire against it.

Developers: Cohort 1 starts 16 November 2026. Companies: every candidate we present has passed all 9.

Apply for Cohort 1Book a call with a founder