Upsky Research

Exploring what AI can do next.

We benchmark and compare models, fine-tune them, build our own, and publish what we learn.

Research areas

How we find out what models can actually do.

01

Benchmarks

We measure models on tasks that reflect real work: reasoning, coding, extraction, and tool use, not just leaderboard favorites.

  • Reasoning
  • Coding
  • Tool use
  • Latency and cost
02

Evals

We design evaluation suites for specific products and domains, so teams can tell whether a change actually made a system better.

  • Eval suites
  • LLM-as-judge
  • Regression testing
  • Human review
03

Fine-tuning

We adapt open models to narrow tasks and measure when fine-tuning beats prompting, retrieval, or simply using a larger model.

  • LoRA
  • SFT
  • Preference tuning
  • Distillation
04

Model comparisons

We put models side by side with the same tasks, prompts, and budgets, so choosing one is a decision backed by data.

  • Quality
  • Cost
  • Speed
  • Open vs. closed
05

Our own models

We are building our own language models to understand the full stack, from data and training to inference and deployment.

In development

  • Language models
  • Data
  • Training
  • Inference

Publications

What we will publish.

Upsky Research is just getting started. As experiments finish, their methods, data, and results will appear here.

First publications in progress

  1. 01

    Research reports

    Long-form write-ups of experiments, with methodology, results, and what we would do differently.

  2. 02

    Leaderboards

    Benchmark results that we update as new models are released.

  3. 03

    Eval suites

    Datasets and test harnesses, shared whenever possible so results can be reproduced.

  4. 04

    Model notes

    Documentation for the models we build: data, training, limitations, and intended use.

  5. 05

    News and analysis

    Short takes on new models, papers, and releases, and what they change for people building software.

How we work

Research that holds up outside the lab.

  • Show the method

    Results come with the context needed to understand them: prompts, settings, and scoring.

  • Test real work

    We prefer tasks taken from real products over puzzles that only exist in benchmarks.

  • Compare fairly

    Same inputs, same budget, and the same scoring for every model we test.

  • Put it to use

    What we learn goes straight into the products and systems we build.

Follow the research.

Write to us to receive new reports, benchmarks, and analysis when they are published.

Have a model, dataset, or question you want us to test? We want to hear it.

Get updates info@upsky.org