In partnership with

Hey {{first_name | there}},

Every time a new model drops, researchers first check how well it does on Math Olympiad problems, science and trivia. 

Then they argue online about whether it's the smartest model in the world. 

But the thing is, nobody hires a vice president based on their SAT scores.

So why do we choose AI models based on benchmarks that have nothing to do with our own actual work?

Official benchmarks are useful for researchers, but they're useless for you. 

They don't tell you which model will write better emails, build better dashboards, or create better presentations for your specific needs.

So I build my own personal benchmarks.

And you can too. It's simpler than you think.

Step 1: Let Claude understand your work

It will first ask you a bunch of questions to get to know your work. Be detailed and specific.

Then based on your work it will give you 5 prompts.

Step 2: Open a side-by-side comparison tool

Now you need a place to test these prompts across multiple models at the same time.

Go to openrouter

It lets you select multiple models, enter a single prompt, and see everything play out in real time. 

The key is: same prompt, same parameters, every model on equal footing. You're comparing the models, not accidentally comparing your settings.

Step 3: Run your prompts and see who wins

Run each prompt through 3-4 models side by side.

Compare a large model like Fable against a small one like Haiku. Compare models from different providers, like Anthropic and OpenAI. Run the same task with the same model three times to see how much it varies.

You'll be surprised. Sometimes the cheap model performs just as well as the expensive one for your specific task. Sometimes a model you ignored turns out to be perfect for your work.

This isn't about finding the "best" model overall. It's about finding the best model for you.

Step 4: Compare price and speed

Once you know which models give you the best output for your tasks, check their price and speed.

Go to Artificial Analysis. It compares AI models across intelligence, pricing, output speed, latency, context window, and more.

The cheapest model isn't always the best value. You need to balance quality, speed, and cost for your specific workload.

And you’re done, here’s how you create your own benchmarks and judge models.

Just know that your benchmark will evolve

When Fable came out, it was so good at open-ended work that I realized I'd been giving AI tasks that were too small. Then my benchmark had to evolve with my more ambitious use of AI.

This is how you "surf the models." Models are getting better every day. To get the most from them, we have to keep up.

AI knows everything that's written down, but using AI models in your work generates judgment they couldn't have been trained on by anyone else. 

Try it yourself. Build your first personal benchmark today.

And if you do, share your results. I'd love to see what you discover.

- Aashish

Is Your Training Data Actually Model-Ready?

If you're fine-tuning a speech model, you've probably hit this wall: DNSMOS gives you a score, but it doesn't tell you whether the data behind that score is actually right for your model. 

Treat it as a pass/fail gate and you'll end up training on audio that looks clean on paper but drags down real-world performance—while good source data gets tossed for no reason.

Voices' CTO DJ Jalali (with the team's senior audio and voice data engineers) just published a free white paper that breaks down the four-step calibration framework they use internally to set model-specific quality thresholds instead of trusting the raw DNSMOS number. It also covers where DNSMOS breaks down and how Voices validates audio for custom datasets at scale.

Reply

Avatar

or to participate