The Rad Journey by Bargava | Substack
The Rad Journey by Bargava
Using AI to make better decisions
By Bargava Subramanian
· Launched 3 months ago
Did ORPO Actually Improve Our $50 Healthcare SLM?
The raw evaluation results: large reasoning gains, uneven specialty performance, and two safety regressions we cannot ignore.
Jun 18•Raghotham Sripadraj
Evals Are Where a Small Language Model Proves It Is Not Lying
How we built an evaluation system that could catch convincing answers, biased judges, and hidden clinical failures
Jun 11•Raghotham Sripadraj
Reinforcement Learning is where SLM learns what it got wrong
How we used the simplest preference optimization technique that could possibly work to teach a 1.7B parameter clinical model its own failure modes
Apr 23•Bargava Subramanian and Raghotham Sripadraj
Supervised Fine-Tuning is where a Small Language Model learns to stop guessing
How we trained a 1.7B parameter clinical reasoning model to think like a Doctor
Apr 14•Bargava Subramanian and Raghotham Sripadraj
Data Curation is the true moat: Here’s how we built a high-signal training data
Most AI teams over-invest in compute and under-invest in data. We did the opposite. And it worked.
Apr 5•Bargava Subramanian and Raghotham Sripadraj
We built a Healthcare SLM for $50 that matched frontier models for our task
Narrow specialized intelligence can be compressed.
Mar 24•Bargava Subramanian and Raghotham Sripadraj
Who am I and why subscribe?
About Bargava
Nov 4, 2019•Bargava Subramanian