Callosum
Visit websiteResearch Engineer, Evals - Member of Technical Staff
Salary not disclosedOnsite
- Engineering
- London
- Full time
- Today
About the role
Callosum is building the infrastructure to unify heterogeneous compute across the full stack for AI systems. This role focuses on developing rigorous, evidence-based evaluation methods for agentic systems to understand performance, localize failures, and guide design decisions. You will design benchmarks and evaluation suites to measure complex agentic behaviors and build the infrastructure that informs the company's future development.
Responsibilities
- Design benchmarks and evaluation suites for agentic behaviour including multi-turn, long-horizon, and tool-using tasks
- Build methodology for statistical power, variance analysis, and construct validity
- Red-team evaluations and design sanity checks to ensure result quality
- Transform raw traces into structured evidence such as failure taxonomies and behavioural signatures
- Develop predictive models to infer system capabilities from partial evidence
- Build durable evaluation and observability infrastructure for company-wide use
Required skills
- LLMs
- Agentic systems
- Statistical analysis
- Experimental design
- Software engineering
Nice to have
- Bayesian inference
- Model selection
- Significance testing
- Uncertainty quantification
- Automated grading
- Model-based judging
Qualifications
- Evidence of independent research capability
- Deep hands-on experience with LLMs in agentic settings
- Strong engineering and coding skills
- Published track record in a relevant field
Benefits
- Competitive Salary
- Equity & Ownership
- Private healthcare
- Visa sponsorship
- Relocation benefits
About the Company
Callosum is the Intelligent Systems Company building the software orchestration layer that co-evolves models, workflows, and silicon into one system. They are based in London and focus on delivering inference tailored to every workload by unifying heterogeneous compute.