llama.cpp-ML is an unofficial downstream fork of ggml-org/llama.cpp focused on native LoRA and QLoRA training for GGUF models without PyTorch.
The project keeps the original llama.cpp inference goal intact and adds a separate ML training direction for local Linux-first experimentation.
Research / technical validation. Not production-ready.
- Linux CPU: correctness validated (Ryzen 9 5950X)
- Linux CUDA: correctness + performance validated (NVIDIA RTX 4090)
- Android S10 CPU: LoRA/QLoRA technical validation PASS (Samsung Galaxy S10)
- Native LoRA/QLoRA training — trainable LoRA A/B tensors in C/C++, frozen GGUF base model
- GGUF adapter save/reload — adapters saved as LoRA GGUF, compatible with existing llama.cpp loader
- CPU + CUDA backends — Q4/QLoRA validated; Q4 ~24% faster than F16 in tested profile
- JSONL / SFT datasets — prompt/completion and messages input, assistant-only label masking
- Quality guard runner — post-epoch validation with holdout prompts
# Build
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j
# Train LoRA adapter (Q4 model, CPU)
./build/bin/llama-finetune \
-m model-q4.gguf \
-f data.jsonl \
--train-max-windows 32 \
-c 256 \
-b 64 -ub 64 \
--lora-train-preset late-attn-qkvo \
--lora-train-rank 4 \
--lora-train-alpha 4 \
-epochs 1 \
--optimizer sgd \
-t 8 -tb 8 \
-o adapter.gguf
# Run inference with the trained adapter
./build/bin/llama-cli \
-m model-q4.gguf \
--lora adapter.gguf \
-p "Hello" \
-n 64- ML Training Details — models, backends, datasets, all test results with numbers
- Android Status — device info, CPU smoke, Adreno/NPU roadmap
- Getting Started — build, config, training, adapter verification
Based on ggml-org/llama.cpp.
Inference works exactly as upstream. Training is a separate addition.
