Stanford
University
  • Stanford Home
  • Maps & Directions
  • Search Stanford
  • Emergency Info
  • Terms of Use
  • Privacy
  • Copyright
  • Trademarks
  • Non-Discrimination
  • Accessibility
© Stanford University.  Stanford, California 94305.
Model Release Update: BMC-CLIP1.1 Scaling Experiment | Stanford HAI
Skip to content
  • About

    • About
    • People
    • Get Involved with HAI
    • Support HAI
    • Subscribe to Email
  • Research

    • Research
    • Marlowe (opens in new tab)
    • Research Programs
    • Grants
    • Student Affinity Groups
    • Centers & Labs
    • Research Publications
    • Research Partners
  • Education

    • Education
    • Executive and Professional Education
    • Government and Policymakers
    • K-12
    • Stanford Students
  • Policy

    • Policy
    • Policy Publications
    • Policymaker Education
    • Student Opportunities
  • AI Index

    • AI Index
    • AI Index Report
    • Global Vibrancy Tool
    • People
  • News
  • Events
  • Industry
  • Centers & Labs

Stay Up To Date

Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.

Sign Up For Latest News

Navigate
  • About
  • Events
  • AI Glossary
  • Careers
  • Search
Participate
  • Get Involved
  • Support HAI
  • Contact Us
news

Model Release Update: BMC-CLIP1.1 Scaling Experiment

Date
June 06, 2025
Topics
Sciences (Social, Health, Biological, Physical)

By Min Sun, Stanford Data Science Scholar—reposted with permission

Earlier this year, we released BMC-CLIP, a CLIP model continually pre-trained on a subset of BIOMEDICA, a dataset comprising 24M image-caption pairsand 31M image-inline text figure reference pairs from scientific literature. Continually pretraining with a batch size of 4,096 (4 × H100) achieved state-of-the-art zero-shot classification and retrieval performance on 40 biomedical tasks using 10× less compute than prior models. 

Today, with support from Marlowe and Stanford Data Science, we’re releasing two new BMC-CLIP-1.1 models trained with larger batch sizes of 8,192 (8 × H100) and 32,768 (32 × H100). To the best of our knowledge, this is the largest biomedical CLIP experiment conducted to date. This is important as larger batches are key for contrastive learning, offering more negative pairs and enhancing the model’s ability to differentiate between relevant and irrelevant image-text matches.

🧪 As expected: Scaling helps. Bigger batch sizes lead to faster convergence and more stable training loss. But interestingly, we also observe checkpoint-specific tradeoffs, some early epochs outperform final ones on specific domains, achieving strong SOTA performance in some tasks! Given these findings, we are releasing all checkpoints for transparency and research.  We hope the community can leverage the full open-source nature of our contributions to understand biomedical domain adaptation and model merging. 

📂 All checkpoints will be made available here: http://bit.ly/4jVFFIl

📂 All models + data: huggingface.co/BIOMEDICA

📄 BIOMEDICA Paper: https://arxiv.org/abs/2501.07171

🙏 Once again, huge thanks to Marlowe and Stanford Data Science for generously providing the compute, these large-scale experiments (8× and 32× H100!) simply wouldn't have been possible without their support. Their infrastructure enabled us to explore how scaling batch size impacts CLIP-style models in biomedicine. We also appreciate the support from NVIDIA to allow us to support these experiments. 

Shout-out to Alejandro Lozano, James Burgess, Serena Yeung-Levy, and Rob Tibshirani for supervising this experiment. 

🚀 The future of multimodal biomedical AI is open source.


This article is a part of the Stanford Data Science legacy publication. Read more about the HAI and Stanford Data Science merger.

Share
Link copied to clipboard!

Related News

Stanford Scientists Build an AI Lab Partner
Nikki Goth Itoi
Jul 09, 2026
News
DNA molecule spiral. 3d rendering

Biomni can analyze mountains of medical data, spot patterns humans might miss, and even design experiments—helping researchers make discoveries faster in the race to cure disease.

News
DNA molecule spiral. 3d rendering

Stanford Scientists Build an AI Lab Partner

Nikki Goth Itoi
Sciences (Social, Health, Biological, Physical)Jul 09

Biomni can analyze mountains of medical data, spot patterns humans might miss, and even design experiments—helping researchers make discoveries faster in the race to cure disease.

How AI Is Accelerating Scientific Discovery
Nikki Goth Itoi
Jul 08, 2026
News

New AI tools generate hypotheses, design experiments, and find patterns in data—transforming how scientists make discoveries across every field.

News

How AI Is Accelerating Scientific Discovery

Nikki Goth Itoi
Sciences (Social, Health, Biological, Physical)Jul 08

New AI tools generate hypotheses, design experiments, and find patterns in data—transforming how scientists make discoveries across every field.

Today's AI Talks Like “Nobody.” New Research Gives It Real Personality.
Jun 08, 2026
News
3D illustration of mirrored human profiles in blue and yellow layers

PsychAdapter lets researchers dial in on personality traits, age, and mental health characteristics to generate text that sounds like real individuals, opening the door to training simulations and personalized content.

News
3D illustration of mirrored human profiles in blue and yellow layers

Today's AI Talks Like “Nobody.” New Research Gives It Real Personality.

HealthcareGenerative AISciences (Social, Health, Biological, Physical)Jun 08

PsychAdapter lets researchers dial in on personality traits, age, and mental health characteristics to generate text that sounds like real individuals, opening the door to training simulations and personalized content.