Substrate Bio
Member of Technical Staff, Intelligence
Salary not disclosedHybrid
- Intelligence
- London
- Full time
- Today
About the role
This role involves building and owning the analysis and ML infrastructure for assay data within a stealth AI and biology startup. You will work closely with protein characterisation scientists and the software team to ensure data quality, traceability, and clear presentation of results. The position requires a blend of data engineering and biostatistics to turn raw instrument output into reliable, versioned datasets.
Responsibilities
- Develop and maintain analysis and ML pipelines that transform assay and sequence data into consistent, versioned datasets.
- Perform data validation to ensure accuracy and identify issues such as mislabelled plates or swapped samples.
- Conduct sequence-level bioinformatics, including computing protein properties and checking constructs.
- Apply biostatistics to experimental data, including experimental design, error propagation, and assay acceptance criteria.
- Implement curve fitting for various assay types with automated QC and confidence measures.
- Detect batch and plate effects and maintain control charts for reference standards.
- Create data visualisations that communicate results, uncertainty, and QC metrics to stakeholders.
Required skills
- Python
- pandas
- NumPy
- SciPy
- SQL
- Biostatistics
- Data modelling
- Data validation
- Experimental design
- Statistical process control
- Data visualisation
- Machine learning pipelines
Nice to have
- R
- Docker
- Kubernetes
- AWS
- LIMS
- Electronic lab notebooks
- Protein language models
- Structure prediction
Qualifications
- Experience with messy experimental data in biotech, pharma, or academic settings
- Equivalent practical experience in a data-heavy field
Benefits
- 30 days annual leave
- Public holidays
- Pension with 10% employer contribution
- Bupa private health cover
About the Company
Substrate Bio is a stealth startup operating at the intersection of AI and biology, focused on building infrastructure for data generation with built-in quality and traceability.