Callosum
Visit websiteResearch Engineer, Benchmarking
Salary not disclosedOnsite
- Engineering
- London
- Full time
- Today
About the role
This role involves building a unified, reproducible benchmarking system to evaluate agentic and algorithmic solutions. You will design a harness that measures task success, quality, and robustness through sandboxed execution rather than self-reported scores. The system you create will serve as the definitive evidence for internal decision-making and external customer proof points.
Responsibilities
- Design and build a unified system for evaluating agentic and algorithmic solutions across various agent topologies and decomposition strategies.
- Mine real commits and traces to build task suites that reflect real agentic work such as code search, edit, and repair.
- Enforce controls against contamination, overfitting, and metric gaming to ensure stable, re-runnable baselines.
- Compare algorithmic and agentic approaches to determine their impact on resolved-task quality, latency, and cost.
- Lead external benchmark co-publications held to rigorous peer and customer review standards.
- Review quality claims across the company and provide data-driven insights to inform product shipping decisions.
Required skills
- Python
- LLM evaluation
- Agentic systems
- Distributed execution
- CI
- Applied AI
Nice to have
- Execution-based grading
- Open-source evaluation
- Code-agent workloads
Qualifications
- PhD in computer science, machine learning, or a related field
- Equivalent research track record
- Authorship or co-authorship of a benchmark or evaluation paper at a recognised venue
Benefits
- Competitive Salary
- Equity & Ownership
- Private healthcare
- Visa sponsorship
- Relocation benefits
About the Company
Callosum is an Intelligent Systems Company building the infrastructure that unifies heterogeneous compute across the full stack. They focus on software orchestration that co-evolves models, workflows, and silicon to deliver inference tailored to specific workloads.