Job Openings Judgment Labs — Full-Stack Engineer

About the job Judgment Labs — Full-Stack Engineer

Judgment Labs — Full-Stack Engineer

Type: Full-time | On-site | San Francisco, CA (5 days/week in-person) Compensation: $200,000–$300,000 base + equity Hiring count: 2 Visa sponsorship: Case-by-case for truly exceptional candidates (H-1B, O-1, OPT); primary scope is candidates who don't require sponsorship Reports to: Founding team (no named contact on role page)

About Judgment Labs

Judgment Labs builds infrastructure for Agent Behavior Monitoring (ABM). Where traditional observability logs exceptions and latency, ABM surfaces behavioral anomalies — instruction drift, context-retrieval loss — in scaled production environments. Hundreds of teams building autonomous agents rely on Judgment to understand how their systems behave post-deployment.

The team has raised $30M+ across two rounds in the past five months, backed by Lightspeed, SV Angel, Valor Equity Partners, Nova Global, Chris Manning, Michael Ovitz, Michael Abbott, Cory Levy, and Kevin Hartz. Under 20 people, shipping at 50+ company velocity, with Olympiad medalists, debate champions, and competitive athletes. Everyone is either an ex-founder or a founder-to-be.

Founded: N/A | Team size: <20 | Total funding: $30M+ Industry: AI agent infrastructure / observability Website: judgmentlabs.ai Office: San Francisco, CA

Why Candidates Should Join

  • Category-defining space: Building the ABM layer beyond traditional observability — how autonomous agents are understood in production.
  • Well-funded, moving fast: $30M+ across two rounds in five months, top-tier backers, under-20 team shipping at 50+ company pace.
  • Real product ownership: Not spec implementation — you talk to customers, define what to build, build it, and iterate end-to-end.
  • Elite, intense team: Olympiad medalists, debate champions, ex-founders; everyone builds like a founder.
  • Comp + perks: Up to $300K (up to $400K for exceptional AI-savvy FDE leads), full benefits, Equinox membership, private chef.

Intake Call Summary

  • No intake call transcript was provided on the role page. An intake video is available on the Contrario listing but was not captured here.
  • Key hiring signals inferred from the page: strong bias toward genuine software-engineering depth (not solutions/FDE-only backgrounds), agent/LLM fluency, and candidates for whom this product-engineering role is a genuine first choice rather than a fallback.

The Role

Own the product experiences and agent infrastructure that make the agent improvement loop legible and actionable for engineering teams. Spans the data layer to the UI — agent swarm interfaces, verification platforms, and the SDK layer that lets developers summon Judgment mid-development. ~30% customer-facing.

What You'll Be Doing

  • Shape how the Judgment Agent runs large-scale parallel investigations across thousands of production traces, merging failure modes, tool errors, regressions, and drift signals into a single actionable answer.
  • Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior.
  • Design how engineers understand long traces, tool calls, decisions, and failures — making a thousand-step reasoning trajectory legible in minutes.
  • Build the swarm UX so engineers can watch parallel investigations, redirect investigators going down the wrong path, and consume findings without reading hundreds of reports.
  • Own the improvement loop: workflows that turn production trajectories into datasets, judges, and regression checks so found-problem to verified-fix feels like one motion.
  • Build and maintain the SDK and terminal-first experience so agent-dev sessions can summon Judgment as a subagent mid-development.
  • Own platform infrastructure: workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments.

Tech stack: Not specified on role page. Full-stack (data layer through UI), SDK development, terminal-first tooling, agent/LLM infrastructure.

Requirements

  • 3 to 7 years full-stack engineering experience
  • Hands-on LLM or agent-building experience
  • End-to-end production system ownership, data layer to UI
  • Customer-facing communication ability (approximately 30% of role)
  • SF in-person 5 days per week, relocation supported

Green Flags

  • Technical background (e.g. coding competitions, research experience) in addition to solutions experience
  • Genuine first choice for a product-engineering role with agent depth
  • Prior evals, observability, or behavior-monitoring product background
  • 4 to 5 years of experience with strong communication track record
  • Strong engineering pedigree from a solid product company
  • Founder background or clear founder-to-be signal

Red Flags

  • FDE is a fallback, not their top choice ("if I can't get X, I'll do FDE")
  • Low commitment signals: slow to book or drops off before the interview
  • Over-indexed on agent experience at the expense of engineering quality
  • Pure solutions or forward-deployed background without real software engineering depth
  • Big-tech candidate using this role as a backup

Role Details

  • Salary — $200,000–$300,000 (junior ~$200K at 3 yrs; mid-to-senior up to $300K at 3–7 yrs; exceptional AI-savvy FDE leads up to $400K case-by-case)
  • Equity — Yes (amount not specified)
  • On-site policy — SF in-person, 5 days/week; relocation supported
  • Visa sponsorship — Case-by-case for exceptional candidates (H-1B, O-1, OPT); main scope is candidates not requiring sponsorship
  • Employment type — Full-time
  • Location — San Francisco, CA

Benefits & perks: Full benefits package · Equinox membership · Private chef · Competitive compensation · Direct customer interface and influence on product roadmap

Screening Questions

None provided on the role page.

Interview Process

Stage 1 — Recruiter screen — Initial screen. Stage 2 — Evals FDE round — Evals / FDE-focused round. Stage 3 — Coding IQ round — Coding / problem-solving round. Stage 4 — FDE onsite — Onsite. Stage 5 — Offer Extended Stage 6 — Candidate Hired — Candidate accepts and starts.

(Platform scheduling states — "Pending Approval," "booked evals round," "booked iq round" — omitted as pipeline artifacts.)

Ideal Companies & Backgrounds

Updated from role page Target product / infra / AI companies — Palantir, Databricks, Datadog, Cognition AI, Decagon, Sierra, Linear, Cursor, Ramp, Figma, Vercel, CockroachDB, Modal Labs, Anyscale, Runway, Applied Intuition, Anduril, Notion, Nomic AI, MotherDuck

Body callouts: solid product companies (e.g. Roblox, Snapchat, Tesla); top-tier infra (e.g. Databricks) is a bonus.

Note: several entries on the page (".Store," "Data Center ISH," "Retoolers," "Mercury Insurance," "SearchHounds," "Ponderosa Agency") appear to be logo-resolver artifacts rather than genuine target companies and are excluded above.

Ideal Candidate Profiles

For reference only — DO NOT CONTACT. No LinkedIn URLs were exposed on the role page. Yuval Danino · Bhagyashri Badgujar · Smriti Sridhar · Elie Harik · Kabeer Thockchom · Sukrit Rao · Sujan Rachuri · Ishan Mehta · Krrish Chawla · Joseph Tey · Aditya Tadimeti · Sathvik Nallamalli · Aliyan Ishfaq