Amgen · 2026–Present
Protein language models & multi-omics infrastructure
Reproducible benchmark and evaluation workflows for AMPLIFY protein-language-model checkpoints, plus Spark SQL and Delta Lake pipelines for multi-omics data.
- Checkpoints benchmarked
- 120M / 350M
- Published context length
- 2,048 residues
Published AMPLIFY comparison: the 350M checkpoint has 43× fewer parameters and 24–29× higher inference throughput than ESM2-15B, depending on sequence length.
- 01IngestMulti-omics data
- 02BenchmarkAMPLIFY checkpoints
- 03EvaluateTask + efficiency metrics
- 04PackageBedrock + Docker
- 05TrackMLflow workflows