Member of Technical Staff, Evaluations
IT
USD 175k-250k / year + Equity
Who we are
Arena Physica is on a mission to accelerate hardware innovation that powers human progress. Our name is inspired by Theodore Roosevelt's 'Citizenship in a Republic' speech. To us, entering the Arena means committing fully and accepting the risk of failure in pursuit of an audacious, worthy cause. We believe the future belongs to those brave enough to build it.
Our team of 60 combines AI engineering and applied physics expertise with deep experience in enterprise deployments. We're headquartered in NYC with presences in San Francisco and Los Angeles, backed by ~$90M from Initialized, Founders Fund, Goldcrest Capital, Fifth Down Capital, and Shield Capital.
If you're ready to do the most important work of your career, join us in the Arena.
What we do
At Arena Physica, we're building electromagnetic superintelligence. Our AI platform Atlas operationalizes physics-grounded intelligence to verify, debug, and optimize hardware across its lifecycle. Atlas is already trusted globally by the world's most advanced hardware companies, including AMD, Anduril, and Bausch & Lomb, for applications across R&D, integration testing, production assembly, and field repair.
About the role
Hardware engineering data lives behind the walls of the organizations building frontier devices. Unlike code, there is no internet-scale dataset of hardware engineering data for frontier models to train on. Much of what we ship to our customers requires consistently asking: where and how are frontier models actually getting better at the engineering work? As our Member of Technical Staff for Evaluations, you will own the means of producing that answer.
You will lead our technical work on evaluating models and harnesses across electrical engineering and the broader set of hardware engineering tasks. You will build the measurement discipline and seamless automated pipelines that report automatically, design the analyses that explain why performance looks the way it does, and craft verifiable/calibrated scoring systems.
The work spans hard infrastructure and deep engineering domain knowledge. In a given week, you might:
- Harden an automated reporting pipeline
- Design an ablation to understand why a harness change caused a regression
- Analyze agent trajectories to find where a model loses efficacy on a dense datasheet and convert the problem into a measurable, tracked phenomenon
- Work with an electrical engineer to build synthetic datasets that push models and harnesses to their limits
- Design and calibrate an LLM judge for a task where ground truth is expensive, and prove it agrees with expert engineers well enough to rely on
The goal is to map the jagged edges of machine intelligence. You will report directly to the CTO.
How you will contribute
- Own the design and reliability of our evaluation harnesses and build automated, regular reporting pipelines. You will make regressions impossible to miss and results legible to engineers, leadership, and our partners
- Design ablations that isolate the effect of individual changes to models, harnesses, tools, and context, and attribute contributors to good/bad performance/cost/runtime
- Build general analyses over agent trajectories that surface why the agent fails. Identify performance clusters and root causes such as context loss, tool inefficiency, engineering gaps, weak spatiotemporal reasoning, and emerging failure modes.
- Leverage agent trajectories to build improvement loops for skills and harnesses across engineering tasks.
- Craft judging and scoring systems that are verifiable wherever possible, and design principled methods for calibrating and validating LLMs used as judges.
- Partner with our Operations and Electrical Engineering teams to hand-craft synthetic datasets that stress the capabilities that matter and cover the long tail of real engineering work.
- Interface with technical counterparts at partner organizations and customers, translating their notions of "good" into measurable artifacts, and communicating what our evaluations show.
- Set the standard for evaluation rigor across the company as capabilities, tools, and domains expand.
Qualifications
- Strong Python skills, including building research or production infrastructure, data pipelines, and tooling that needs to be reliable at scale.
- Demonstrated experience designing evaluations, benchmarks, or metrics for ML systems, ideally for large language models and/or agentic systems.
- Hands-on experience with LLM agents: harnesses, tool use, and analysis of agent trajectories.
- Strong grounding in experimental design and statistics, and sound judgment about what makes a metric trustworthy versus misleading.
- Excellent written and verbal communication, especially when explaining technical results to specialists and non-specialists alike.
- A bias toward automation, you'd rather build the pipeline once than run the analysis by hand ten times.
- [Preferred] BBackground or working fluency in electrical or hardware engineering (e.g. circuit design, RF/EM, signal integrity, PCB layout, test and measurement), or a track record of ramping quickly on a technical domain and reasoning about it rigorously.
- [Preferred] Proficiency in calibrating and validating LLM-as-judge systems.
- [Preferred] Experience sourcing, curating, or hand-crafting datasets, particularly synthetic datasets built to probe specific capabilities.
- [Preferred] Expertise in evaluating multimodal, spatial, or spatiotemporal reasoning.
- [Preferred] A track record of building dashboards and reporting that people actually trust and use.
- [Preferred] Experience working directly with customers or external technical partners.
Benefits & Perks Include:
- 100% of the monthly premiums covered with Aetna medical vision, and dental insurance for you and your dependents
- 401(k) Retirement Plan
- Unlimited PTO
- Lunch every day from local restaurants via Sharebite
- Relocation support provided
The base salary range for this position is $175,000 - $250,000 yr. However, base pay offered may vary depending on job-related knowledge, skills, and experience. In addition to base salary, we also offer competitive equity and benefits packages.
This position may require access to information protected under U.S. export control laws and regulations, including the Export Administration Regulations (EAR) and the International Traffic in Arms Regulations (ITAR). Please note that any offer for employment may be conditioned on authorization to receive software or technology controlled under these U.S. export control laws and regulations without sponsorship for an export license.