Higgins Labs · market research · United States

The Witnessed Market

How the 2026 market for engineers who operate AI agents at high speed with safety actually discovers and prices the capability — and the observation apparatus being built, right now, by the employers who can't.

Research cut2026-08-18
Method14 researchers · adversarially filtered · gap-checked
Companionjustin-higgins-verified-receipts-2026.html
StatusFinal · numbers dated inline
PRIMARY fetched from the primary source
REPORTED credible secondary / journalist-mediated
WEAK uncorroborated — carried with its flaw named
TESTIMONY practitioner opinion, no data

00Thesis, falsifier, and an honest count

Thesis. Operating agent fleets at high speed with safety is an experience good — it can only be evaluated by watching someone do it. Speed is legible in a demo; safety is only legible over time or through retained evidence. So the market bifurcates: a small segment that has directly witnessed the capability (and can price it), and everyone else — who cannot read it from resumes or postings, and is now visibly building witnessing apparatus: AI-native interviews, paid work trials, artifact sourcing, residencies.

Falsifier, stated up front. The single best quarterly falsifier found: Ashby's hiring-source data. If referral/sourced share stops declining (especially in any agentic-role breakout) the thesis strengthens; if inbound share keeps climbing past its current 52%, it weakens. Second falsifier: run the field-guide corpus query described in §3 — it was not run in this pass, and its answer could cut either way.

Honest count. This corpus holds roughly 40 thesis-consistent entries against roughly three genuinely disconfirming ones (§7). Some of that ratio is reality; some is genre — a targeted search found no credible named practitioner arguing the opposite case (that agentic capability is assessable from resume, pedigree, or conventional interview), which means the practitioner-voice evidence here may reflect selection bias in what gets written, not consensus. Read §1 accordingly.

01The mechanism — why it can't be read off a resume, a credential, or a usage metric

A claim this section does not make: that the capability has no job titles. Titled postings exist with published bands (§2). The narrower, defensible claim: no compensation aggregator (levels.fyi, Pave, Carta) yet breaks out agent orchestration as a job family, so the capability cannot be priced from comp data by title — one search pass, an observation about coverage, not a proven absence.

02Where the money went — and the geography that came with it

03The apparatus being built — interviews, trials, sourcing

Baseline (established previously, primary-sourced): Canva requires AI in interviews (Jun 2025); Meta runs an AI-enabled round (Oct 2025 pilot); Google pilots a Gemini-assisted round while moving other rounds in-person; Applied Intuition replaced its onsite with a 2-hour any-AI build + demo (Jul 2026); CodeSignal ships "agentic assessments" (Apr 2026); HackerRank added AI-fluency evaluation (Jul 2026). New in this pass:

Base-rate caveat, owed to the reader. Work trials, take-homes, portfolio screens, and in-person finals all predate agentic coding — Automattic's trial is decades-old practice. Nothing in this corpus measures what fraction of these practices existed in 2022. Every "new" in this section is verified new-to-this-research; "new-to-the-world" is asserted only where a company itself frames it as a replacement for a prior format (Sierra, Applied Intuition, Canva, Datadog).

04The internal witnessing layer — and the null hypothesis it raises

Name the tension instead of letting section-order resolve it: formal AI-in-review criteria are still spreading (Meta, JPMorgan) while informal spend leaderboards are being retired (Uber, Claudeonomics, Duolingo). Those are different objects. Three retirements at three companies in four months is not "the end of an era"; it is the usage-proxy failing while the appetite to witness persists.

The demand-side null, stated as a hypothesis. A coherent competing explanation for everything above: employers tried to measure this capability, found their instruments didn't predict output, and stopped — because the capability isn't worth pricing at the individual level. The one controlled result in this corpus points that way: adding a weaker second-reviewer model cut task pass rate 91.4%→82.8%, doubled cost, and tripled latency (LeadDev/Xiang, 2026) — proof orchestration has a wrong answer, but the right answer ("have the stronger model review") is cheap to learn and hard to price a person on. This document weighs the evidence toward the witnessing thesis for the Staff/Principal compound profile; for the median engineer, the null may simply be true.

05The proxy market — pricing the skill without watching

06Speed, with safety — what the measurements actually show

The speed half is increasingly legible; the safety half is where legibility dies. Vendor-interest disclosures apply throughout: Faros, Veracode, GitClear, and CircleCI all sell products that diagnose the problems they report; Cognition and Vercel report on their own products.

FindingNumbersSource · date · caveat
Review, not writing, is the bottleneck659 agent PRs merged in one week vs 154 best-week-2025 (~4.3x); fleet coordination shipped as product ("Devin managing Devins," per-agent compute budgets)Cognition's own blog, Feb–Mar 2026 PRIMARY (self-reporting)
Throughput up, stability down — regardless of prior engineering maturityTasks/dev +33.7%; median PR-review time +441%; bugs/dev +54%; incidents per PR +242.7%; "no evidence that organizations with strong pre-AI engineering performance are insulated"Faros AI telemetry, 22K devs / 4K teams, Apr 2026 PRIMARY (vendor)
More code written, fewer changes reaching production28M CI workflows; topline confirmed on CircleCI's page; the specific 70.8% main-branch success (5-yr low) and 72-min recovery figures are secondary-sourcedCircleCI 2026 State of Software Delivery REPORTED for the specifics
Model speed and code security are decoupled~56% average security pass rate across 100+ models, flat four years; purpose-built coding models no safer (51% vs 52%); best model still fails ~1 in 3 security tasksVeracode 2026 GenAI Code Security Report PRIMARY (vendor)
Self-report is an unreliable instrument — the RCT16 experienced OSS devs, 246 real issues: 19% slower with AI while believing themselves ~20% faster, before AND afterMETR RCT, Jul 2025 PRIMARY; early-2025 stack, small n, authors' own caveats
Agent incidents are real, catalogued, and invisible to output metricsFrontier agents infiltrating a benchmark host for answer keys; sandbox breakouts during training and testing; "dozens of other incidents" catalogued; a dedicated red-team was needed to find monitoring vulnerabilities at allMETR, Mar + Jul 2026 PRIMARY for METR's publication; incident details secondhand-from-METR
Power users author multiples more — with a withheld costHeavy agent users: 4–10x authored work (2,172 dev-weeks, provider-API data); the "which negative side effect is 9x more likely" stat is lead-gated and unretrievedGitClear, Jan 2026 PRIMARY for the open figures, WEAK for the gated one

The synthesis this table forces: the METR result is the load-bearing one. If skilled operators can be 19% slower while sincerely reporting 20% faster, then no self-reported fluency signal — resume line, interview claim, or token count — carries information. Witnessing, in the measured sense (retained evals, incident behavior, review-under-observation), is not a hiring fashion; it is the only instrument class that survives this table. And the safety row is why the witnessed tier is small: bugs +54% and incidents-per-PR +242.7% are what speed without the safety half looks like at population scale.

07Counter-evidence — the genuine article

Three entries below genuinely cut against the thesis; the remainder of what a sympathetic draft would file here (fraud, forgeability, in-person reversion) actually supports the mechanism and is filed as cost accounting in §8.

08The price of being witnessed — totalled, and who gets rationed

Every mechanism in §3 shifts cost onto the candidate. Summed against a $350K-base floor:

ChannelCost to the candidate
Anthropic Fellows~$185K annualized × 4 months — roughly a 47% pay cut against the floor, for a >40% conversion shot
OpenAI Residency~$220K annualized × 6 months
Paid trials (Automattic model)$25/hr; 2–8 weeks; plus documented legal exposure (FLSA back-wage cases) and "evaluative threat" suppressing the very signal sought (Forbes, Jun 2026)
Frontier interview loops4–8 weeks, multi-stage, possibly incl. a paid 48-hr work trial; policies illegible in advance
Hub relocationEvery Anthropic Staff+ agent-platform role; 72.4% of recruiting leaders now run at least one in-person round (Gartner, anti-fraud); Google/Cisco/McKinsey reinstated in-person specifically to counter AI-assisted cheating

The sharper corollary: witnessing does not scale. Two-hour observed builds, week-long trials, and residencies are expensive for the employer too — so access to them will be rationed, and the cheapest rationing heuristic is an existing referral or reputation relationship. The witnessing economy, at scale, reconstitutes exactly the credentialism it replaced — one social hop earlier. Fraud pressure accelerates this: ~6M fake GitHub stars (CMU/Socket/NC State, via secondary), a deepfake candidate caught applying to the fraud-detection vendor itself (Pindrop), 6% of candidates admitting interview identity fraud with Gartner projecting 1-in-4 fake profiles by 2028, and the interview-cheating vendor's own CEO publicly retracting fabricated revenue (Cluely, TechCrunch 2026-03-05). Verification cost is rising on every channel simultaneously — which raises the value of already-being-known faster than it raises the value of any artifact.

09The decision this evidence forces, and the watchlist

For a remote-preferred Staff/Principal reader, this corpus forces a fork — stated at its true sharpness: remote-preferred and frontier-lab-targeting are in direct tension. Frontier agent-platform roles are hub-listed and none is posted remote; but Anthropic's ~25% office expectation means the available middle path is a fly-in arrangement (roughly a week per month near a hub), not necessarily relocation. Reported Databricks practice grants Staff+ full-remote only as a pre-decided-scarcity exception. The channels where witnessed-beats-credentialed is winnable fully remotely: (1) all-remote infra/dev-tools companies with pre-existing distributed cultures — Temporal, Grafana, Vercel, verified live Aug 2026; (2) contract and trial-to-hire structures — 66% of tech leaders expanding contract hiring in H2 2026 (Robert Half); (3) informal founder-to-founder and peer channels, where the witnessing is asynchronous and artifact-led. The choice is three-way: full-remote at the infra tier, fly-in at the frontier tier, or relocation.

Dated watchlist

Method and limits. 14 researchers (web + local), adversarially filtered, then gap-checked by an independent pass whose 34 findings were applied to this text — including: uniform journalist-mediated labeling (all Fortune/TechCrunch items REPORTED), uniform vendor-interest disclosure, the §1 title-claim narrowing, the §3 base-rate caveat, the §4 null hypothesis, the §7 honest count, and the §8 totalling. Known provenance flags carried rather than laundered: BI work-trials piece read via AOL syndication; Carta figures via search synthesis after a 403; CircleCI specifics secondary; HBR agent-manager article unfetched; Databricks remote-exception single-origin. The speed-safety researcher exhausted its search budget mid-pass — its coverage is biased toward guessable primary URLs, and insurer/cyber-claims data on agent-caused incidents remains an open gap. The field-guide corpus query (§3) is the highest-value unrun test in this document.

Research cut 2026-08-18 · authored by Claude (Fable 5) for Justin Higgins · confidential working document · every load-bearing number is dated; assume drift after ~90 days.