Skip to content
Machine Learning Tutorial

Fine-Tuning an Open Model on a Single GPU

LoRA changed the economics of fine-tuning. You can adapt a 7B model to your style of writing on a single consumer GPU in a few hours.

P Priya Sharma Updated 5 min read
Fine-Tuning an Open Model on a Single GPU

Fine-tuning used to sound like a data-centre project. Then came LoRA — low-rank adaptation — which freezes the original model and trains only a tiny set of extra parameters. A 7 billion parameter model can be adapted on a single consumer GPU in hours.

When to fine-tune (and when not to)

Fine-tuning changes style and behaviour, not knowledge. Use it to teach a model your tone, your output format, or a specialised task like summarising legal clauses. It cannot reliably inject new facts — that is what RAG is for.

If you only need a few examples, try few-shot prompting first. Fine-tuning pays off when you have hundreds or thousands of examples and a consistent task.

Prepare the data

Collect pairs of input and ideal output. Quality beats quantity: 500 clean, consistent examples usually beat 5,000 noisy ones. Format matters more than size.

[{"instruction": "Summarize the meeting notes",
  "input": "...",
  "output": "..."}]

Consistency is the secret. If some outputs are bullet points and others are prose, the model will mix them. Decide the format first, then write every example in it.

Choose a base model

Pick an open-weight model in the 7B–14B range: LLaMA, Mistral or Qwen families are solid starting points. Smaller models fine-tune faster and deploy on less hardware; larger models generalise better on harder tasks. Start small and scale only if quality demands it.

Configure LoRA training

With the peft library, LoRA configuration is a few lines:

from peft import LoraConfig, get_peft_model

config = LoraConfig(
    r=16,            # rank — bigger captures more, costs more
    lora_alpha=32,
    target_modules=["q_proj", "v_proj"],
    lora_dropout=0.05,
)

model = get_peft_model(base_model, config)
model.print_trainable_parameters()  # typically ~1-2% of weights

Only 1–2% of parameters are trainable. That is the entire trick: the model keeps its learned knowledge, and you add a lightweight adaptation layer on top.

Train on one GPU

Use 4-bit quantisation to fit the base model into memory, batch size 1–2 with gradient accumulation, and a low learning rate around 2e-4. Watch the loss on a held-out validation split. Two to four epochs is a common sweet spot; more epochs can overfit your style until it sounds like a parody of itself.

Evaluate, then deploy

Hold back 10% of examples the model never trained on. Compare its outputs against your ideal outputs on a few dimensions: format compliance, tone, factual accuracy. Run the fine-tuned model side by side with the base model on the same prompts.

When quality looks right, merge the LoRA weights into the base model and export. You now have a compact, deployable model that writes in your voice.

Fine-tuning with LoRA is the closest thing to a free lunch in applied AI: a few hours of training, a tiny model delta, and a dramatically more consistent assistant.

A quick recap

To summarise, this guide is organised around the core ideas below, and each one matters for a different reason.

  • When to fine-tune (and when not to) — Fine-tuning changes style and behaviour, not knowledge.
  • Prepare the data — Collect pairs of input and ideal output.
  • Choose a base model — Pick an open-weight model in the 7B–14B range: LLaMA, Mistral or Qwen families are solid starting points.
  • Configure LoRA training — With the peft library, LoRA configuration is a few lines:
  • Train on one GPU — Use 4-bit quantisation to fit the base model into memory, batch size 1–2 with gradient accumulation, and a low learning rate around 2e-4.

Questions worth asking yourself

Use these prompts to turn the article into decisions about your own setup.

  • How does when to fine-tune (and when not to) apply to the way you approach fine-tuning on one GPU today?
  • How does prepare the data apply to the way you approach fine-tuning on one GPU today?
  • How does choose a base model apply to the way you approach fine-tuning on one GPU today?
  • How does configure LoRA training apply to the way you approach fine-tuning on one GPU today?

Putting it into practice

Applying fine-tuning on one GPU is less about memorising every feature and more about building a repeatable routine. Start with the single task that costs you the most time each week, run it through the workflow described above, and keep a short note of what changed. Your own results are a better guide than any generic benchmark. The same principles show up wherever you work with LLaMA, Fine-tuning, Python.

The Machine Learning landscape moves quickly, so treat what you have read as a starting point rather than a fixed rulebook. Revisit the tools and techniques you rely on every few months, retire anything that no longer earns its place, and fold in only the additions that solve a problem you actually have.

Further reading

If this Machine Learning topic was useful, these related guides go deeper on the areas you are most likely to need next.

P

Written by

Priya Sharma

Priya previously built ML systems at a cloud provider. She writes hands-on tutorials covering embeddings, RAG and model deployment.

More articles by Priya Sharma →

Frequently asked questions

How long does it take to read this article?

Most readers finish in under ten minutes. Use the table of contents to jump to the section you need.

Do I need previous experience to follow along?

No. We explain every concept as it appears, and the code examples are self-contained.

Report an issue with this page

Comments

Leave a comment

Comments are moderated and will appear once approved.