All solutions

Models and AI engineering

A general model knows everything except how your company works.

It has read the internet and none of your tickets. We take an open-weight model, train it on the work your team has already done, and measure it against the one you pay per token for today. When ours wins, it runs in your own infrastructure, costs a fraction per request, and the weights are yours.

2 weeks
to a measured baseline that says whether training is worth it at all
Your weights
the trained model is yours, with nothing tying it to us or to a vendor
10–40×
the usual drop in cost per request once a small model takes over a high-volume job
ES · EN
a senior team in your time zone, in both languages

What we train

Four kinds of model work, and most companies need one.

We start from the task, not from the technique. Sometimes the honest answer is that you do not need a trained model at all, and we would rather say so before anyone signs anything.

Where we usually start

Fine-tuned open models

We take an open-weight model, keep training it on the work your company has already done, and end up with one that answers and reasons the way your team does. It runs on your hardware and the weights are yours.

  • Adaptation of Llama, Qwen, Mistral, and the rest
  • LoRA or full fine-tunes, whichever the job needs
  • Trained on your tickets, contracts, and past decisions
  • Benchmarked against the API model you use today
Where we usually start

Purpose-built task models

Most of what companies ask a large model to do is one narrow job repeated a million times: sort this, pull that field, score this call. A small model trained for that one job is faster, cheaper, and far easier to defend in an audit.

  • Classifiers for routing, triage, and tagging
  • Extractors that pull fields out of documents
  • Rankers and recommenders over your own catalog
  • Forecasting and anomaly detection on your history

Reinforcement learning and preference tuning

When the right answer is a matter of judgment rather than a label, we train on the judgment itself: what your reviewers accepted, what they rewrote, and what they threw out.

  • Reward models built from your reviewers' decisions
  • Preference tuning on real corrections, not synthetic pairs
  • Policies for agents that have to act, not just answer
  • Feedback loops that keep improving after launch

Evaluation and serving infrastructure

The part that decides whether any of the above survives contact with production: a test set that catches regressions, a registry that says which version is live, and serving that holds up under your real traffic.

  • Eval harnesses built from your own hard cases
  • Quantization and distillation to fit your budget
  • GPU or CPU serving, autoscaled or on-premise
  • Monitoring for drift, cost, and quality

Before you train anything

Training is the fourth thing to try, not the first.

Every step down this list costs more than the one above it and takes longer to change your mind about. We work down it in order, and most engagements stop before the bottom.

  1. 01

    Prompting a frontier model

    Typical effort: Days

    A good prompt against Claude or GPT with your context pasted in. Cheap to try and cheap to abandon, which is exactly why it goes first.

    Worth it when

    The task is rare, varied, or changes every month.

    Not enough when

    You run it a million times a day and the invoice shows it.

  2. 02

    Retrieval over your own content

    Typical effort: Weeks

    The model stays general and you hand it your documents at question time. Most “the model does not know our business” problems are really this problem.

    Worth it when

    The gap is knowledge the model never had.

    Not enough when

    The gap is behavior: format, tone, or judgment it keeps getting wrong.

  3. 03

    Fine-tuning an open model

    Typical effort: 1–2 months

    You teach the model how the work is done by showing it the work being done. This is where most of our model engagements land.

    Worth it when

    You have thousands of examples of the task done right.

    Not enough when

    Nobody can agree yet on what “done right” looks like.

  4. 04

    Reinforcement learning

    Typical effort: 2–4 months

    You train on judgment instead of labels: what got accepted, what got corrected, and what got rejected outright.

    Worth it when

    Quality is a matter of preference and you have reviewers.

    Not enough when

    A plain label would have captured the same thing.

  5. 05

    Training from scratch

    Typical effort: Rarely

    A model built from the ground up. For a handful of narrow, well-defined tasks this is genuinely the right answer. For anything resembling a chatbot, it almost never is.

    Worth it when

    The task is narrow, the data is yours, and nothing pretrained fits.

    Not enough when

    You want a chatbot. Adapt one instead.

We charge for the evaluation that answers this question, and we charge for it separately, so hearing “you do not need to train anything” costs you two weeks instead of a project.

Where it pays off

The jobs that eat an afternoon, every afternoon.

None of these need a model that can write poetry. They need one that is right about your work, fast, and cheap enough to run on everything that comes in.

  • Ticket and email routing

    Every message sorted by team, urgency, and language before a person opens it, trained on ten years of how your team actually sorted them.

  • Document extraction

    Invoices, delivery notes, and contracts turned into fields your systems can read, including the scanned ones and the note somebody wrote in the margin.

  • Quality and defect detection

    A model trained on your own images or sensor history that flags the bad batch on the line instead of at the customer.

  • Demand and risk forecasting

    Trained on your sales, your seasons, and your disruptions, rather than on a generic curve that has never seen your market.

  • Call and conversation scoring

    Every call scored against the criteria your own supervisors use, so coaching covers all of them instead of the three someone had time to listen to.

  • Compliance and policy review

    A first pass over documents against your internal rules, with the clause it tripped on attached, so your reviewers start from a shortlist.

How we engineer it

What separates a demo from a model you can run a business on.

A notebook that scored well once is not a deliverable. These are the things we build alongside the model, and together they are usually more work than the training itself.

  • Evaluation first

    We write the test set before we train anything. If we cannot measure the task, we cannot tell you whether we improved it.

  • Reproducible runs

    Every run is pinned: data version, code version, seed, and hardware. A result nobody can reproduce is a rumor.

  • Data lineage and consent

    We record where every training example came from and what you are allowed to do with it, before it ever reaches a GPU.

  • Versioning and rollback

    Models are versioned like code. You can see which one answered a given request, and go back to the previous one in minutes.

  • Cost and latency budgets

    A target for price per request and response time is set at the start. A model that misses it is not finished, however good the accuracy looks.

  • Drift monitoring

    Your business changes and the model quietly stops being right about it. We watch for that and tell you when it is time to retrain.

How we work

From “could a model do this?” to one in production.

The first phase exists to talk you out of it. It is cheaper for both of us to find out in two weeks that a task is not learnable than in four months.

  1. 01Estimate: 2 weeks

    Feasibility and baseline

    We define the task precisely, build the test set, and measure what an off-the-shelf model already scores on it. That number is what everything afterwards has to beat.

    DeliverableA test set, a baseline score, and a recommendation

  2. 02Estimate: 2–4 weeks

    Data curation

    We assemble the training set out of work your team has already done, clean it, label what needs labeling, and hold back the hard cases for the exam.

    DeliverableA versioned training and evaluation set

  3. 03Estimate: 3–6 weeks

    Training

    We run the experiments, compare sizes and techniques, and stop at the smallest model that clears the bar you set.

    DeliverableTrained weights and the runs that produced them

  4. 04Estimate: 1–2 weeks

    Evaluation and review

    We score it against the baseline and put its worst outputs in front of the people who do the job today, because a benchmark never catches everything.

    DeliverableAn evaluation report and a go or no-go

  5. 05Estimate: 2–3 weeks

    Deployment

    We serve it inside your infrastructure behind an API your systems can call, with a fallback for the cases where it is not sure.

    DeliverableA served model with monitoring and runbooks

  6. 06Estimate: Ongoing

    Retraining and operation

    New examples pile up, quality drifts, and the model gets retrained on a schedule your team can run without us in the room.

    DeliverableA retraining pipeline and the handover

The timings on each stage are a reference from past projects.

Stack

What we train and serve with.

We use what you already pay for wherever it holds up. When there is no constraint, this is where we start.

Base models
  • Llama
  • Qwen
  • Mistral
  • Gemma
  • DeBERTa
Training
  • PyTorch
  • Hugging Face
  • LoRA / PEFT
  • TRL
  • Axolotl
  • Unsloth
Classical ML
  • scikit-learn
  • XGBoost
  • LightGBM
  • Prophet
Evaluation
  • Custom harnesses
  • Weights & Biases
  • MLflow
  • Ragas
Serving
  • vLLM
  • TGI
  • ONNX Runtime
  • Triton
  • llama.cpp
Infrastructure
  • AWS
  • GCP
  • Modal
  • Docker
  • Terraform
  • Kubernetes

How we engage

Three ways in, depending on how sure you are.

Most companies start at the top and decide after it, with a number in hand instead of a hunch.

01

Feasibility study

Two to three weeks, fixed price, to answer one question: can a model do this well enough to be worth building, and what would it cost to run?

Best for

Teams with a task in mind and no idea whether it is learnable.

Includes

  • A precisely defined task and a test set
  • A baseline from off-the-shelf models
  • A cost and latency projection
  • A written recommendation, including “do not build this”
02

Model build

End to end: data, training, evaluation, and a served model your systems can call, with your team in the room for all of it.

Best for

A task that cleared feasibility and has a budget behind it.

Includes

  • Data curation and labeling
  • Training and experiment tracking
  • Evaluation against your baseline
  • Deployment inside your infrastructure
  • Weights, code, and runbooks handed over
03

Model operations

A monthly retainer for the part nobody plans for: the model is live, the world moved, and something needs retraining.

Best for

Teams running models in production without an ML engineer on staff.

Includes

  • Scheduled retraining and evaluation
  • Drift and cost monitoring
  • Serving upgrades as base models improve
  • An engineer who knows your models on call

Your model

You own the weights, not a license to them.

Everything we train runs in your accounts, on base models whose licenses let you keep what comes out. When the engagement ends you hold the weights, the training data, and the code that produced both, and you have no technical reason to call us again.

  • Training and serving in your own cloud accounts
  • Open-weight base models with licenses that permit commercial use
  • Your data never trains anything outside your project
  • Full handover: weights, datasets, pipelines, and runbooks

Questions

What companies ask us first.

How do we know whether we even need a custom model?

Often you do not, and the feasibility study is how we prove it either way. If a prompt and your own documents get you to 90% of the value, that is the answer we will give you, and it costs two weeks to find out instead of a project.

How much data do we need?

For a classifier or an extractor, a few hundred good examples is often enough to beat a general model. For fine-tuning a language model to reason the way your team does, a few thousand. Quality matters far more than count, and “we have ten years of tickets” is usually a better starting point than people expect.

What if our data is a mess?

It is. Everyone's is. Cleaning and labeling is a real phase with real weeks attached, which is why it sits on the plan above instead of hiding inside the training step. The one thing we cannot fix is data that never recorded the thing you want to predict.

What does this cost, and what does it cost to run?

A feasibility study is fixed price and a small fraction of a build. A full build depends on the task, and training is rarely the expensive part: data and evaluation usually are. Running it is the easier number, because a small task model on your own hardware is generally much cheaper per request than an API call, and we project that for you before you commit.

Why not just use Claude or GPT for everything?

For plenty of things you should, and we build those systems too. Training earns its place when volume makes per-token pricing hurt, when latency matters, when the data cannot leave your building, or when the task is narrow enough that a small model beats a large one at it. If none of those apply, we will point you at an API and save you the project.

Do we need to buy GPUs?

For training, no. We rent them for the weeks we need them and pass the cost through. For serving it depends on the model: plenty of task models run fine on CPUs you already have, and the ones that need a GPU usually need less of one than people assume.

Tell us the job your team does a thousand times a month.

Describe one repetitive judgment call somebody on your team makes all day. We will tell you whether a model could learn it, what it would take to find out properly, and what it would cost to run once it works.

The first conversation is free. If a prompt would solve it, we will tell you that instead of selling you a training run.

info@upsky.org