Naveen Vivek

ProjectsInference

Local fine-tuning and inference

QLoRA fine-tuning of Qwen2.5-Coder-7B on a home RTX 5070 Ti to fix common code-quality issues, plus local inference with Ollama.

Year
2026
Area
Inference
Built with
Unsloth, QLoRA, Qwen2.5-Coder-7B, Ollama, RTX 5070 Ti
Code
github.com/naveenvivek/finetuningunsloth

A home-lab project for learning how models really behave on limited hardware: teaching a small open model one narrow skill on a single consumer GPU with 16 GB of memory.

I fine-tuned Qwen2.5-Coder-7B with QLoRA to fix common code-quality issues, and run local inference through Ollama to see where memory, quantization and throughput set the limits. A 12 GB quantized model runs at about 70 tokens per second.