Supervised Fine-Tuning
Train on curated input and output pairs to lock in response format, domain vocabulary, tone, and structured output that parses reliably on the first attempt.
Baaz fine-tunes open-weight models so a small model you own performs like a frontier model on your task, at a fraction of the cost per token. We work in parameter-efficient tuning, LoRA and 4-bit QLoRA on open weights from 7B to 70B, engineered for deterministic output your code can parse every time. Because the weights are self-hosted, training and serving can run in your own data centre or fully air-gapped, which is often the only route when sensitive data cannot reach a hosted API. And because we modernise the systems these models feed, the tuned model lands in a working application instead of a notebook.
Fine-tuning is worth doing when a model keeps getting the shape of the answer wrong: the wrong format, the wrong vocabulary, the wrong level of caution, the wrong reasoning style. Prompting and retrieval fix a lot of that, and we will tell you when they are enough. Past that point the fastest route is to train the behaviour into the weights. We come at this as a software delivery problem. The measure of success is a model running inside your product under load, returning output your code can parse every time.
Each of these ships with an evaluation set, a held-out regression suite, and a documented recipe you can rerun without us.
Train on curated input and output pairs to lock in response format, domain vocabulary, tone, and structured output that parses reliably on the first attempt.
Parameter-efficient training that updates a small set of weights instead of the whole model. For larger open weights we run 4-bit QLoRA with NF4 quantisation, double quantisation, and paged optimizers, which is what brings 7B to 70B models within reach of sensible hardware.
Training sets built against a strict JSON or SQL schema, with deliberate abstention examples so the model returns a clean null for an attribute it cannot find instead of inventing a plausible one. This is usually what separates a demo from something you can put behind an API.
DPO, ORPO, KTO, and GRPO for the behaviours supervised examples cannot express: which of two acceptable answers is better, when to refuse, when to ask instead of guessing.
Move the behaviour of a large model into a model in the 1B to 8B range that you own and serve. For high-volume structured work this is where the cost curve changes by an order of magnitude.
Mining real production failures, deduplicating, decontaminating against your test set, and generating synthetic coverage for gaps. A few thousand well-scored examples beat tens of thousands of scraped ones.
A golden dataset built from real failures, deterministic checks first and a calibrated judge second, wired into CI so a regression blocks the release instead of reaching users.
We build and modernise ERPs, CRMs, logistics engines, and computer-vision stacks, so the tuned model goes into an actual workflow with the error handling, retries, fallbacks, and audit trail that implies.
The frameworks, methods, and serving targets we use to take an open-weight model from a base checkpoint to a monitored production endpoint.
Families we tune and benchmark against each other, from 7B task models up to 70B.
Single-GPU speed for iteration, distributed training when the run outgrows one card.
The 4-bit path that makes large open weights trainable without a cluster.
Preference and reward optimisation for the behaviours examples cannot state.
Sharded training across multiple GPUs, with long-context runs where the task needs them.
Throughput-oriented serving with per-request adapter selection and quantised weights.
The layer that keeps a malformed response from reaching your database.
The measurement layer that decides whether a checkpoint ships.
Fine-tuning is the wrong tool about as often as it is the right one. These are the signals we look for before recommending it, and if none of them hold we will say so.
Fine-tuning reliably changes how a model answers. It does not reliably teach it new facts. If your gap is knowledge that changes weekly, retrieval is the fix and we will build that instead.
A response that is right nine times in ten is a bug when the tenth goes into a database. Schema-strict tuning with abstention is the reliable route to parseable output and an honest null.
The best training set is the log of what your current model got wrong in production. Teams with that history get a useful model in weeks. Teams without it need instrumentation first.
At high request volume on a narrow task, a tuned small model changes the unit economics and the tail latency. At low volume, a frontier API is cheaper than the GPUs and the engineering.
Plenty of AI work dies between a promising notebook and a shipped feature. If that is where yours is, the missing piece is usually engineering, not more modelling.
Some data simply cannot be sent to a hosted API, and no contract makes that acceptable: patient records under HIPAA, card and transaction data under PCI, personal data with residency conditions attached, defence and critical-infrastructure work, or the drawings, formulas, and source code that are the business itself. A hosted frontier model means every prompt crosses your boundary to a third party who logs it, and an internal security review will stop that. A tuned open-weight model is the way out, because the weights run on hardware you control, inside your own network or fully air-gapped, and the sensitive text never leaves the building.
Hosted models get versioned, deprecated, and quietly changed underneath you, which is a problem when a regulator asks you to reproduce a decision from eighteen months ago. A model you hold as a file does not move unless you move it.
We agree the pass mark first and build the harness to measure it. Training without an eval you trust produces a model nobody can approve.
We work through your production traces to separate prompting problems, retrieval problems, and genuine behaviour problems. Only the third kind needs a fine-tune.
We pin down the exact JSON or SQL contract the model has to satisfy, including how it should signal that a value is genuinely absent.
Curate, score, and deduplicate real examples, write the abstention cases explicitly, generate synthetic coverage for the gaps, and hold back a slice that training never sees.
We benchmark open-weight candidates from 7B to 70B against your eval, because which base wins changes from task to task and only the eval can tell you.
LoRA or 4-bit QLoRA runs with tracked hyperparameters, checkpoints, and cost per run, so any result can be reproduced or rolled back later.
Where judgement matters, a preference pass over ranked outputs teaches the model which acceptable answer is the better one.
The candidate runs against the held-out suite plus general-capability slices, so we catch a model that learned your task and forgot how to reason.
The model goes behind a service in your actual application, with schema validation, retries, fallbacks, and logging, so a bad response degrades instead of corrupting data.
Deploy behind a versioned endpoint with adapter hot-swapping, then keep scoring live traffic so quality drift surfaces before your users report it.
We have built software since 2018 from Bengaluru with a US office in Sheridan, Wyoming, and we treat fine-tuning as production engineering. The training run is the short part. Everything around it decides whether the model is safe to ship.
The harness exists before the first training job. You get a number you can defend internally, not a demo that looks better than it measures.
If prompting, retrieval, or a better base model solves your problem, that is the recommendation. Fine-tuning you did not need is the most expensive kind.
Strict schemas and explicit abstention examples are part of the dataset from day one, so a missing attribute comes back as a null your code can handle instead of a confident guess.
We modernise ERPs, CRMs, logistics engines, and vision stacks for a living, so the model lands inside a real workflow with the plumbing that requires.
Training runs in your cloud account or ours under your terms, with open weights, no lock-in to a hosted tuning service, and an air-gapped path when you need one.
You get the weights, the dataset, the recipe, and the eval harness. Retraining next quarter should not require another statement of work.
Most work starts small on purpose. A pilot answers whether tuning helps your task at all before anyone commits to a pipeline, and the eval harness it produces is what the larger engagement is judged against.
One task, one base model, one adapter. We build the eval harness, assemble a first dataset from your traces, train, and report the measured gap against your current baseline. The deliverable is a decision backed by numbers, plus everything needed to go further.
Dataset pipeline, schema and abstention design, adapter and preference training, regression gates in CI, serving with hot-swappable adapters, integration into the application that consumes it, and drift monitoring after launch. Scoped per task, per model size, and by how much of your existing system has to change.
Engagements in between are common and scoped the same way. If your problem turns out not to need fine-tuning, the pilot is where we tell you.
A fine-tuned open-weight model is an artifact you control, so your constraints decide where it runs. We serve it wherever the data is allowed to be, and keep the base model shared so several adapters can run behind one endpoint.
LLM fine tuning continues training an existing language model on your own examples so its behaviour changes: the format it answers in, the vocabulary it uses, the reasoning style it follows, and when it declines. The result is a model you own that is specialised to your task, rather than a general model you steer with a long prompt on every request.
Usually both, in that order. Retrieval is the right tool when the gap is knowledge, especially knowledge that changes, because you can update a document store instantly and fine-tuning does not reliably teach new facts. Fine-tuning is the right tool when the gap is behaviour that a prompt keeps failing to enforce. The sequence we recommend is prompt, then retrieval, then fine-tune, then distil, stopping as soon as the eval is satisfied.
By training for abstention on purpose. The dataset is built against a strict JSON or SQL schema and includes examples where an attribute is genuinely absent and the correct answer is a null, so returning nothing becomes a learned behaviour instead of a failure mode. On top of that, schema validation and constrained decoding sit in front of your application, so a malformed response is caught before it reaches your data.
With a scoped validation pilot: one task, one base model, one adapter. We build the evaluation harness first, assemble a dataset from your own traces, train, and report the measured gap against your current baseline. That gives you a decision backed by numbers before anyone commits to a full pipeline, and the harness carries forward into the larger engagement if you go ahead.
Less than most teams expect, if the examples are good. A few hundred to a few thousand carefully curated and deduplicated examples routinely outperform tens of thousands of scraped ones, because the scraped set teaches inconsistency. What matters is that the examples come from real failures, follow one schema, and are scored for correctness before they go in.
Open-weight families, mainly Llama, Mistral, Qwen, Gemma, and Phi, from around 7B up to 70B. Parameter-efficient tuning is what makes that range practical: LoRA for light task adapters, and 4-bit QLoRA using NF4 quantisation, double quantisation, and paged optimizers when the model is too large to train in full precision. We benchmark several candidates against your eval instead of defaulting to a favourite.
At sustained volume on a narrow task, generally yes, and often by around an order of magnitude per token once you account for a small tuned model on your own GPUs instead of a large rented one. At low or spiky volume the arithmetic reverses, because you pay for idle capacity and for the engineering to run it. We model this against your actual request pattern before recommending either way.
It can, and that failure is called catastrophic forgetting. A model trained hard on one narrow task can lose general reasoning and instruction following while scoring well on the thing you measured. We guard against it by keeping general-capability slices in the regression suite, so a candidate that gained on your task and lost elsewhere fails the gate before it ships.
Yes, and for a lot of enterprises that is the whole reason to fine-tune an open-weight model. Both training and serving can run on hardware inside your own data centre, including fully air-gapped with no outbound connectivity, so prompts containing regulated or commercially sensitive data never cross your network boundary. That is usually the deciding factor under HIPAA, PCI, or data-residency obligations, and in defence, critical infrastructure, and any setting where the input itself is the intellectual property. The alternative we also support is your own VPC on AWS, GCP, or Azure, which satisfies most residency requirements while leaving operations to your cloud team.
Yes. Training runs in your own cloud account or on your hardware, and because the base models are open weight there is no requirement to send examples to a hosted tuning service that may retain them. The dataset, the weights, the recipe, and the evaluation harness are all yours at the end, so nothing about the result depends on continued access to us or to a model vendor.
With a harness built before training. We agree the pass criteria up front, hold back a slice of data that training never sees, run deterministic checks first and a judge calibrated against human-labelled examples second, and report the comparison against your current baseline. A checkpoint that does not clear the bar does not ship, and you get the evidence either way.
Foundational pre-training research is not what we do, and neither is chasing benchmark results for their own sake. Our work starts where a model has to become working software. Annotation-only projects go to Payana, our managed data-annotation service.
Ready to scope this stack? Brief the Baaz squad or browse more services.