Benchmarks
We measure models on tasks that reflect real work: reasoning, coding, extraction, and tool use, not just leaderboard favorites.
Upsky Research
We benchmark and compare models, fine-tune them, build our own, and publish what we learn.
Research areas
We measure models on tasks that reflect real work: reasoning, coding, extraction, and tool use, not just leaderboard favorites.
We design evaluation suites for specific products and domains, so teams can tell whether a change actually made a system better.
We adapt open models to narrow tasks and measure when fine-tuning beats prompting, retrieval, or simply using a larger model.
We put models side by side with the same tasks, prompts, and budgets, so choosing one is a decision backed by data.
We are building our own language models to understand the full stack, from data and training to inference and deployment.
In development
Publications
Upsky Research is just getting started. As experiments finish, their methods, data, and results will appear here.
First publications in progress
Long-form write-ups of experiments, with methodology, results, and what we would do differently.
Benchmark results that we update as new models are released.
Datasets and test harnesses, shared whenever possible so results can be reproduced.
Documentation for the models we build: data, training, limitations, and intended use.
Short takes on new models, papers, and releases, and what they change for people building software.
How we work
Results come with the context needed to understand them: prompts, settings, and scoring.
We prefer tasks taken from real products over puzzles that only exist in benchmarks.
Same inputs, same budget, and the same scoring for every model we test.
What we learn goes straight into the products and systems we build.
Write to us to receive new reports, benchmarks, and analysis when they are published.
Have a model, dataset, or question you want us to test? We want to hear it.