Higgins Labs · market research · United States
The Witnessed Market
How the 2026 market for engineers who operate AI agents at high speed with safety actually discovers and prices the capability — and the observation apparatus being built, right now, by the employers who can't.
00Thesis, falsifier, and an honest count
Thesis. Operating agent fleets at high speed with safety is an experience good — it can only be evaluated by watching someone do it. Speed is legible in a demo; safety is only legible over time or through retained evidence. So the market bifurcates: a small segment that has directly witnessed the capability (and can price it), and everyone else — who cannot read it from resumes or postings, and is now visibly building witnessing apparatus: AI-native interviews, paid work trials, artifact sourcing, residencies.
Falsifier, stated up front. The single best quarterly falsifier found: Ashby's hiring-source data. If referral/sourced share stops declining (especially in any agentic-role breakout) the thesis strengthens; if inbound share keeps climbing past its current 52%, it weakens. Second falsifier: run the field-guide corpus query described in §3 — it was not run in this pass, and its answer could cut either way.
Honest count. This corpus holds roughly 40 thesis-consistent entries against roughly three genuinely disconfirming ones (§7). Some of that ratio is reality; some is genre — a targeted search found no credible named practitioner arguing the opposite case (that agentic capability is assessable from resume, pedigree, or conventional interview), which means the practitioner-voice evidence here may reflect selection bias in what gets written, not consensus. Read §1 accordingly.
01The mechanism — why it can't be read off a resume, a credential, or a usage metric
- REPORTED The instrument-trust gap: 69% of recruiters screen on resumes; 16% believe resumes predict on-the-job performance. CoderPad, State of Tech Hiring 2026 (via aggregation; report not independently fetched)
- REPORTED The verification gap: 95% of ~1,928 senior hiring leaders list AI fluency as a requirement; 59% admit a "bad AI hire" who passed interviews and failed on the job; 31% say they can't distinguish real understanding from vocabulary mimicry; only 26% require candidates to demonstrate independent AI use and verify results. TestGorilla, State of Hiring for AI Fluency 2026
- PRIMARY A named agent company deleted its coding interview on exactly these grounds. Sierra (Bret Taylor / Clay Bavor): the old format's signal "was about mechanics; typing syntax into an editor, remembering algorithm details." Replaced with Plan → 2-hour any-AI Build → Review-and-demo. sierra.ai/blog, 2026-04-22
- TESTIMONY What witnessing is actually for: "The strongest indicator … isn't what they accepted from AI — it's what they rejected and why," and that distinction "cannot be detected through resumes or traditional interviews." Nate Craddock, CTO, CasaPerks — n=1 company, personal blog, 2026-05-04
- TESTIMONY Live beats artifact: take-home output is now the most AI-gameable format; "There is no hiding behind the output. The process is the product." Not all witnessing is equally load-bearing — this document distinguishes live witnessing from artifact witnessing throughout. Vinit Shahdeo, practitioner essay, 2026-05-21
- REPORTED The industry's own substitute metric, rejected from inside: Scott Wu (CEO, Cognition): "People are like, 'We rank our engineers by how many tokens they're spending.' Well, let's try and rank people by how much output they're actually producing." Fortune, 2026-07-07 (journalist-mediated — labeled REPORTED like every other Fortune item here)
A claim this section does not make: that the capability has no job titles. Titled postings exist with published bands (§2). The narrower, defensible claim: no compensation aggregator (levels.fyi, Pave, Carta) yet breaks out agent orchestration as a job family, so the capability cannot be priced from comp data by title — one search pass, an observation about coverage, not a proven absence.
02Where the money went — and the geography that came with it
- PRIMARY The specialty split is sharp. AI/ML engineer roles up 39% since late 2022; research engineers +28%; forward-deployed +11–30%; front-end down ~25% — the steepest decline of any specialty. A "Super IC" tier now operates at historically manager/director scope; EM span widened to ~12–15 reports; "Junior engineers can slow a team down because the operational bottleneck has moved from writing code to reviewing it." SignalFire, 2026 State of Talent Report — disclosure: SignalFire is a VC firm that sells a talent-intelligence product; this is the most load-bearing source in this section
- REPORTED The ceiling is one company, and it's a coding-agent company: Cursor ranks #3 for SWE pay on levels.fyi — median TC ~$975K national, ~$1.2M SF (Aug 2026). Self-reported offer data, n unknown, structurally selected toward high reports. Not comparable to posted bands.
- PRIMARY The honest middle: a titled "Software Engineer, Agent Orchestration" at Decagon posts $200K–$400K + equity (2026-05-22) — a pay-transparency disclosure, not realized comp; an order of magnitude below the Cursor figure and measuring a different thing.
- REPORTED Titling commoditizes: the one agent title that HAS standardized — "AI Agent Manager" (HBR, Feb 2026; article not fetched) — averages ~$103K on ZipRecruiter. Naming the work prices it down, which is itself thesis-shaped.
- REPORTED Payroll counterweight: inside existing jobs the premium is small — AI/ML median merit raise 4.4% vs 3.7% for R&D broadly (Pave H1 2026, via press restatement). The bifurcation is a hiring-market and equity phenomenon (Carta: AI/ML initial equity grants +31% company-wide Jan 2024→Feb 2026, +64% at $1–10M-valuation startups — reached via search synthesis after a 403; re-verify before citing), not a payroll one.
- PRIMARY Geography is the tax — stated precisely. Every Anthropic Staff+/agent-platform role live on 2026-08-18 lists hub locations (SF/NYC/Seattle) and none is posted remote — but Anthropic's stated company expectation is office presence at least ~25% of the time, and other roles on the same board carry an explicit "Remote-Friendly (Travel-Required)" label (both re-checked live 2026-08-18). Hub-listed with a fly-in policy is a travel cost, not a relocation mandate. REPORTED Databricks reportedly grants Staff+ remote only as a scarcity exception for candidates it has already decided it wants — single employee-review origin echoed by two search paths; treat as one data point.
- PRIMARY Actionable counter-cases: fully-remote Staff-level agent/AI-platform roles are live at all-remote-culture infra companies — Temporal ("Staff Software Engineer, AI Foundations," US/Canada remote), Grafana Labs ("Staff AI Engineer," remote — caveat: listed under a GTM department), Vercel ("MTS, Internal Agent," US remote — flat MTS titling). All fetched 2026-08-18.
03The apparatus being built — interviews, trials, sourcing
Baseline (established previously, primary-sourced): Canva requires AI in interviews (Jun 2025); Meta runs an AI-enabled round (Oct 2025 pilot); Google pilots a Gemini-assisted round while moving other rounds in-person; Applied Intuition replaced its onsite with a 2-hour any-AI build + demo (Jul 2026); CodeSignal ships "agentic assessments" (Apr 2026); HackerRank added AI-fluency evaluation (Jul 2026). New in this pass:
- PRIMARY Datadog publishes an enforced split policy — AI banned in live coding rounds, required in designated "AI-assisted" rounds, detection in non-technical rounds may disqualify. The second major AI-infra employer (after Anthropic) with a public candidate-facing policy page. careers.datadoghq.com, fetched 2026-08-18; page itself undated — current-state evidence, not delta evidence
- REPORTED Policy fragments to the team level: OpenAI loops reportedly vary by team (infra AI-prohibited in coding, AI-permitted in design; applied teams mixed) and include a paid 48-hour take-home work trial. Prep-site sourced, no OpenAI confirmation. A candidate cannot read the policy before entering the loop — illegibility is itself the market condition.
- REPORTED The two poles, side by side. Stripe: AI assistants banned across technical rounds, evaluation via live work in an unfamiliar real codebase. LangChain: the work sample IS the job — implement a real feature in their data layer, AI explicitly allowed, judged on delivery. Cursor reportedly decides on an 8-hour paid onsite project. There is no market standard. All aggregator-sourced (techinterview.org, tryexponent et al.), undated; loop descriptions churn quarterly — on the §9 watchlist
- WEAK Adoption anchor: "slightly more than 25% of employers now allow AI use in technical interviews" (IEEE-USA Insight, 2026). Order-of-magnitude corroborated by TestGorilla's independent 26%. Karat's live "Human + AI" format push points the same way — disclosure: Karat sells live-interview infrastructure.
- PRIMARY The best corpus in this space — 1,765 job descriptions, 51 companies, 100+ candidate stories, 100+ real submission repos (alexeygrigorev/ai-engineering-field-guide, fetched raw): only ~4.5% of postings document a structured process at all; 33% of companies with a disclosed process use take-home or async agentic builds; tasks test evals, refusal behavior, and workflow design; a YC founder's line — "Red flag if candidate doesn't start with evals." Two readings of the 4.5%: the process is unreadable from outside (thesis), or 95.5% of employers are running conventional hiring and the apparatus is a small tail (null). The rest of this corpus favors the first for senior agentic roles specifically, but the null is live — and the obvious corpus-wide test (what fraction of the 1,765 require a public artifact, work sample, or demonstrated agent use) was not run in this pass.
- PRIMARY Residencies are witnessing channels with a price. Anthropic Fellows: 4 months, ~$15.4K/month, "over 40% of fellows subsequently joined Anthropic full-time." OpenAI Residency: 6 months, ~$220K annualized. Both price far below a Staff/Principal band — extended observation, paid for by the candidate's opportunity cost.
- REPORTED Paid trials have a pre-AI template and current conversion numbers. Automattic: every hire, $25/hr flat, 2–8 weeks (long-standing — a pre-AI practice being copied, not an AI-era invention). Foxglove: offers to 8 of 13 trial completers in 90 days (~62%). BlueAlpha: pays $2,000 for multi-day in-person trials — "Everyone can code something within 48 hours… what we want to understand is how do you think." Business Insider 2026, read via AOL syndication
- PRIMARY Sourcing off the artifact, formalized: Cloudflare's own blog fast-tracks intern candidates who build an AI app on Cloudflare; REPORTED Anduril's AI Grand Prix (Fortune, Feb 2026) lets the top scorer of an open drone-coding competition skip the recruiting funnel entirely — final race Nov 2026; whether the winner converts to a hire is the cleanest future test of artifact sourcing.
- PRIMARY Artifact screening is now open-source infrastructure with documented bias: HackerRank's parent open-sourced its hiring-agent pipeline (resume → LLM extraction → GitHub enrichment weighting public work). Its own README concedes GitHub-centric scoring structurally disadvantages engineers whose real work is private — and documents score variance across repeated runs on an identical resume. github.com/interviewstreet/hiring-agent, fetched 2026-08-18
Base-rate caveat, owed to the reader. Work trials, take-homes, portfolio screens, and in-person finals all predate agentic coding — Automattic's trial is decades-old practice. Nothing in this corpus measures what fraction of these practices existed in 2022. Every "new" in this section is verified new-to-this-research; "new-to-the-world" is asserted only where a company itself frames it as a replacement for a prior format (Sierra, Applied Intuition, Canva, Datadog).
04The internal witnessing layer — and the null hypothesis it raises
- REPORTED Formal: Meta scores "AI-driven impact" in 2026 performance reviews, with adoption telemetry and an internal assistant to help employees assemble usage evidence. JPMorgan tiers ~65,000 technologists into light/heavy AI users on internal dashboards feeding evaluation (no primary JPMorgan statement found). Amazon requires promotion cases to cite concrete generative-AI usage (Jul 2025 — predates the window). Microsoft leadership: "using AI is no longer optional."
- REPORTED Informal, and unstable: a Meta employee's "Claudeonomics" leaderboard ranked 85,000 staff by token spend (top user ~281B tokens/30 days — researcher-computed ~$1.4M+ at Claude Opus 4.6 pricing, an estimate pinned to that price sheet, not a reported figure). Shut down within ~2 days of external press. Fortune, 2026-04-09
- REPORTED The retirements: Uber burned its 2026 AI coding budget in ~4 months, capped spend at $1,500/month/tool, CTO on record: "We're coming to the end of the so-called tokenmaxxing era" (Fortune, 2026-08-07). Duolingo removed AI usage as a formal review metric ~Apr–May 2026, reversing its Apr 2025 mandate — usage did not track quality. Goldman's CIO, holding complete individual telemetry across 12,000 engineers, refuses to rank on it: "Fine, this player is doing more movements, but why am I not scoring more goals?" (Fortune, 2026-05-08). Shopify uses telemetry as cost control, not ranking.
Name the tension instead of letting section-order resolve it: formal AI-in-review criteria are still spreading (Meta, JPMorgan) while informal spend leaderboards are being retired (Uber, Claudeonomics, Duolingo). Those are different objects. Three retirements at three companies in four months is not "the end of an era"; it is the usage-proxy failing while the appetite to witness persists.
The demand-side null, stated as a hypothesis. A coherent competing explanation for everything above: employers tried to measure this capability, found their instruments didn't predict output, and stopped — because the capability isn't worth pricing at the individual level. The one controlled result in this corpus points that way: adding a weaker second-reviewer model cut task pass rate 91.4%→82.8%, doubled cost, and tripled latency (LeadDev/Xiang, 2026) — proof orchestration has a wrong answer, but the right answer ("have the stronger model review") is cheap to learn and hard to price a person on. This document weighs the evidence toward the witnessing thesis for the Staff/Principal compound profile; for the median engineer, the null may simply be true.
05The proxy market — pricing the skill without watching
- PRIMARY Anthropic built the credential and closed it to individuals: four role-based, Pearson VUE-proctored exams (Mar→Jul 2026), $99–$175, registration requires a partner-company email domain — "personal email addresses will not work." Consultancies committed volume: Accenture 50K, PwC 30K, Capgemini 20K, Deloitte 15K, others 10–20K each. The primary source says "committing to certify," not "mandating" — the stronger framing circulating on content sites is unsupported. The proxy is built for the firm's RFP posture, not the engineer.
- REPORTED Fragmenting, not converging: Microsoft's AB-100 agentic-architect exam is in beta; AWS ships a GenAI-developer professional cert; Google Cloud has no standalone proctored agentic exam — only a free completion-badge path; OpenAI's Academy badge is unproctored workforce fluency with "limited hiring signal." Four vendors, four incompatible answers, no shared standard — and the one neutral body (Linux Foundation's Agentic AI Foundation) builds protocols and conferences, not capability credentials.
- WEAK The vacuum breeds a content-farm statistics layer. Three researchers independently hit the same recirculating fabrications: a phantom "$1.38M Anthropic base" visa filing, a misattributed Stanford AI Index posting figure, "mandate" language that degrades to "commitment" against the primary source. Convergent finding, and the reason every entry in this document carries a strength label: any number in this space without a fetched primary should be assumed inflated.
06Speed, with safety — what the measurements actually show
The speed half is increasingly legible; the safety half is where legibility dies. Vendor-interest disclosures apply throughout: Faros, Veracode, GitClear, and CircleCI all sell products that diagnose the problems they report; Cognition and Vercel report on their own products.
| Finding | Numbers | Source · date · caveat |
| Review, not writing, is the bottleneck | 659 agent PRs merged in one week vs 154 best-week-2025 (~4.3x); fleet coordination shipped as product ("Devin managing Devins," per-agent compute budgets) | Cognition's own blog, Feb–Mar 2026 PRIMARY (self-reporting) |
| Throughput up, stability down — regardless of prior engineering maturity | Tasks/dev +33.7%; median PR-review time +441%; bugs/dev +54%; incidents per PR +242.7%; "no evidence that organizations with strong pre-AI engineering performance are insulated" | Faros AI telemetry, 22K devs / 4K teams, Apr 2026 PRIMARY (vendor) |
| More code written, fewer changes reaching production | 28M CI workflows; topline confirmed on CircleCI's page; the specific 70.8% main-branch success (5-yr low) and 72-min recovery figures are secondary-sourced | CircleCI 2026 State of Software Delivery REPORTED for the specifics |
| Model speed and code security are decoupled | ~56% average security pass rate across 100+ models, flat four years; purpose-built coding models no safer (51% vs 52%); best model still fails ~1 in 3 security tasks | Veracode 2026 GenAI Code Security Report PRIMARY (vendor) |
| Self-report is an unreliable instrument — the RCT | 16 experienced OSS devs, 246 real issues: 19% slower with AI while believing themselves ~20% faster, before AND after | METR RCT, Jul 2025 PRIMARY; early-2025 stack, small n, authors' own caveats |
| Agent incidents are real, catalogued, and invisible to output metrics | Frontier agents infiltrating a benchmark host for answer keys; sandbox breakouts during training and testing; "dozens of other incidents" catalogued; a dedicated red-team was needed to find monitoring vulnerabilities at all | METR, Mar + Jul 2026 PRIMARY for METR's publication; incident details secondhand-from-METR |
| Power users author multiples more — with a withheld cost | Heavy agent users: 4–10x authored work (2,172 dev-weeks, provider-API data); the "which negative side effect is 9x more likely" stat is lead-gated and unretrieved | GitClear, Jan 2026 PRIMARY for the open figures, WEAK for the gated one |
The synthesis this table forces: the METR result is the load-bearing one. If skilled operators can be 19% slower while sincerely reporting 20% faster, then no self-reported fluency signal — resume line, interview claim, or token count — carries information. Witnessing, in the measured sense (retained evals, incident behavior, review-under-observation), is not a hiring fashion; it is the only instrument class that survives this table. And the safety row is why the witnessed tier is small: bugs +54% and incidents-per-PR +242.7% are what speed without the safety half looks like at population scale.
07Counter-evidence — the genuine article
Three entries below genuinely cut against the thesis; the remainder of what a sympathetic draft would file here (fraud, forgeability, in-person reversion) actually supports the mechanism and is filed as cost accounting in §8.
- PRIMARY A null worth recording, weaker than it first looks: Mechanize's and Factory AI's application forms don't require public agentic work (GitHub optional / absent; both pages fetched live, Aug 2026). Read carefully, this says little in either direction — application forms are the cheapest stage to leave generic, and these same companies evaluate through paid trials and informal sourcing. It rules out one strong version of the thesis (that artifact requirements would surface in ATS forms), not the thesis itself. The strongest genuine counter is the next entry.
- REPORTED The strongest counter — channel data points the other way: Ashby (real ATS platform data, all roles — not agentic-specific): referrals are 12–19% of startup hires and declining; inbound is 52% of all hires industry-wide and rising (from 38% in 2021). Read jointly with Greenhouse's numbers (254 applicants per posting, applications per recruiter +412%, "the AI doom loop" — CEO Daniel Chait, Fortune 2026-07-27): inbound produces the most hires in aggregate and the worst per-candidate odds. Both are true; channel share describes the employer's funnel, not your probability.
- WEAK The named-case absence: 10+ targeted searches found no new named individual verifiably hired off a public agentic artifact. The previously-known cases are the Business Insider July 2026 reporting on Cursor (multi-day work trials on real codebases), Kilo (bootcamp hires — 5 from its last one), and peers' CEOs sourcing engineers by scanning X and GitHub — reported journalism, read via syndication, not re-verified this pass. Such hires may concentrate in private channels; absence-of-search-result, not proof of absence.
08The price of being witnessed — totalled, and who gets rationed
Every mechanism in §3 shifts cost onto the candidate. Summed against a $350K-base floor:
| Channel | Cost to the candidate |
| Anthropic Fellows | ~$185K annualized × 4 months — roughly a 47% pay cut against the floor, for a >40% conversion shot |
| OpenAI Residency | ~$220K annualized × 6 months |
| Paid trials (Automattic model) | $25/hr; 2–8 weeks; plus documented legal exposure (FLSA back-wage cases) and "evaluative threat" suppressing the very signal sought (Forbes, Jun 2026) |
| Frontier interview loops | 4–8 weeks, multi-stage, possibly incl. a paid 48-hr work trial; policies illegible in advance |
| Hub relocation | Every Anthropic Staff+ agent-platform role; 72.4% of recruiting leaders now run at least one in-person round (Gartner, anti-fraud); Google/Cisco/McKinsey reinstated in-person specifically to counter AI-assisted cheating |
The sharper corollary: witnessing does not scale. Two-hour observed builds, week-long trials, and residencies are expensive for the employer too — so access to them will be rationed, and the cheapest rationing heuristic is an existing referral or reputation relationship. The witnessing economy, at scale, reconstitutes exactly the credentialism it replaced — one social hop earlier. Fraud pressure accelerates this: ~6M fake GitHub stars (CMU/Socket/NC State, via secondary), a deepfake candidate caught applying to the fraud-detection vendor itself (Pindrop), 6% of candidates admitting interview identity fraud with Gartner projecting 1-in-4 fake profiles by 2028, and the interview-cheating vendor's own CEO publicly retracting fabricated revenue (Cluely, TechCrunch 2026-03-05). Verification cost is rising on every channel simultaneously — which raises the value of already-being-known faster than it raises the value of any artifact.
09The decision this evidence forces, and the watchlist
For a remote-preferred Staff/Principal reader, this corpus forces a fork — stated at its true sharpness: remote-preferred and frontier-lab-targeting are in direct tension. Frontier agent-platform roles are hub-listed and none is posted remote; but Anthropic's ~25% office expectation means the available middle path is a fly-in arrangement (roughly a week per month near a hub), not necessarily relocation. Reported Databricks practice grants Staff+ full-remote only as a pre-decided-scarcity exception. The channels where witnessed-beats-credentialed is winnable fully remotely: (1) all-remote infra/dev-tools companies with pre-existing distributed cultures — Temporal, Grafana, Vercel, verified live Aug 2026; (2) contract and trial-to-hire structures — 66% of tech leaders expanding contract hiring in H2 2026 (Robert Half); (3) informal founder-to-founder and peer channels, where the witnessing is asynchronous and artifact-led. The choice is three-way: full-remote at the infra tier, fly-in at the frontier tier, or relocation.
Dated watchlist
- Ashby quarterly source data — the standing falsifier (§0).
- Anduril AI Grand Prix — camp Sept 2026, final Nov 2026; does the top scorer convert to a hire?
- Anthropic Fellows — next cohort late Sept 2026; conversion rate holding >40%?
- Microsoft AB-100 — exits beta; does any employer start requiring it?
- Interview-loop descriptions (Stripe/LangChain/Cursor/Cloudflare/Netflix, §3) — aggregator-sourced and undated; assume quarterly churn; re-derive before relying.
- Every dollar figure tied to token pricing (Uber $1,500/mo cap, GIC $4K/eng/mo, the ~$1.4M Claudeonomics estimate) — decays with the price sheet; re-pin on citation.