Sovereign AI · trained from scratch

A language model that never leaves the building.

Kalorad trains its own model — Ukrainian, Russian, English, code — and runs it inside your network. No vendor licence. No external calls. Weights, tokenizer and the learning loop are yours.

KaloradLM 683M · dense decoder · own tokenizer · first full pretraining run completed Sept 2026 on one AMD MI300X

Kalorad · Chaton-prem
you
Що таке суверенна модель? Три речення.
KaloradLM 683M
Message Kalorad…↵
Processes · mi300x-01R→U→T→O→R
pretrain-1b24 GPU 0
96 %
model-server :8000
41 req/min
embedder memory index
312 facts/s
cycle-manager stage T · promote in
06:12:44
KaloradLM · 683M · neural activitylayer 17 / 24
synapses 1 920firing 1 240 / sheads 24 · kv 4
context 4 096precision bf16 · fp32 mastersexternal calls 0
Generating · next-token distributionT = 0.7
pretrain.log · MI300Xlive
Memory · sqlite-vec4 ms
recallWhat did we decide about data residency?
every message · every fact · local · yours
683Mparameters 13.1Btokens 43 hone MI300X $86compute source: pretrain.log · 16 750 steps × 786 432 tokens
Under the hood

Written from scratch. Every layer is ours.

No forked checkpoint, no borrowed tokenizer. The architecture below is implemented in PyTorch in this repository, tested on CPU and verified numerically on ROCm.

Architecture
Dense decoder
hidden 1536 · 24 layers · GQA 24 : 4 heads · SwiGLU 4096
Positions
RoPE + YaRN
context extension without retraining; RMSNorm pre-norm
Optimizer
Muon + AdamW
Newton–Schulz orthogonalised updates for matrices, AdamW for embeddings and head
Schedule
WSD + curriculum
warmup – stable – decay; data mix shifts by phase
Corpus · 12.5B tokens · own SentencePiece tokenizer
en 35 · uk 25 · ru 22 · code 18
EnglishUkrainianRussianCode
Precision
fp32 masters
bf16 autocast for compute; masters stay fp32 so small updates are not lost
Verification
3.3e-06
CPU vs GPU numerical parity on fixed input; 9 training bugs found on hardware, each now a test
Live inside Kalorad

Processes, memory, learning — visible.

An assistant that remembers everything and learns from everything needs an operator’s view. This is what the node is doing right now.

Processes
Memory
Training
node mi300x-01uptime 00:00:00cycle R→U→T→O→R
pidprocessstategpu / cpurate
4112pretrain-1b24kalorad_model/pretrain.py · step 16 750training
88 041 tok/s
4188model-serverOpenAI-compatible · :8000serving
41 req/min
3970corpus-builder14 workers · SentencePiecetokenizing
1.2M tok/s
4201embeddernomic-embed-text · memory indexindexing
312 facts/s
4230memory-compactorsqlite-vec · nightlyscheduled
02:14:09
4007cycle-managerR→U→T→O→R · promotion checkstage T
06:12:44

GPU 0 · utilisation

Self-improvement cycle

Rrun
Uupdate
Ttrain
Ooptimise
Rretire
VRAM178 / 192 GBpower612 Wcheckpointstep_0016750loss4.8819
recallWhat did we decide about data residency?
index sqlite-vec · embeddings nomic-embed-text · local
latency 4 ms · corpus every message, every fact · retention yours

First run · what the curve says

Steps 0 – 13 750: loss flat at 6.1. The dataset shifted labels once more than the model did — for 37 hours it learned to predict the token after next. Found in the curve, fixed on the node, now guarded by test_training_objective.py.
Steps 13 750 – 16 750: 6.1 → 4.88. Same weights, correct objective. Perplexity 132. The next run starts with the fix, the reference Muon and fp32 masters.
tokens13.1Bthroughput88 041 tok/swall43 hcost$86
Console preview. Training figures are from the run log; process rates are illustrative.kalorad_model/reports/run_700m_20260910
Sovereign by construction

Not a policy. A perimeter.

Fine-tunes of Western open models inherit a licence and a dependency. A model trained from scratch inherits nothing — it can live in a closed network for a bank, a ministry or a hospital.

CLIENT PERIMETER · CLOSED NETWORK Employees Documents · DB KaloradLMweights · memory · loop Vendor API0 calls
01

No vendor licence

The weights are trained by us, on our corpus, with our code. There is no upstream terms-of-use to inherit, revoke or re-negotiate.

02

No external calls

Inference, memory and retraining run on the client’s hardware. Nothing a user writes leaves the perimeter — by architecture, not by contract.

03

A model that keeps learning inside

The R→U→T→O→R cycle retrains on the organisation’s own feedback, promotes a new version when it measures better and retires the old one — all inside.

Proof, not promises

One full run. Every number from the log.

10–12 September 2026, a single AMD MI300X on credits. Not a product checkpoint — a proof that the stack trains end to end, and a costed path to the production model.

0
parameters
0
tokens seen
0
optimizer steps
0
wall clock, one GPU
0
total compute
0
tokens / second
0
final perplexity
3.3e-06
CPU / GPU parity
Next: 1.24B × 25B tokens ≈ 163 GPU-hours ≈ $300 at market MI300X prices.kalorad_model/reports/run_700m_20260910/pretrain.log
Roadmap

What compute buys next.

The code is written; each step below is a run, not a research question. Order is fixed by what unblocks the next.

Now

1.24B production run

25B trilingual tokens, corrected objective, reference Muon, fp32 masters.

+1 week

SFT for chat

Short supervised pass for the chat format and refusals.

+2 weeks

Ukrainian benchmark

A published evaluation set — the number a buyer can check.

+1 month

8 vertical LoRA

Legal, medical, finance, public sector… measured against the base.

Later

MoE 3.16B

1.84B active parameters — capacity without the inference bill.

Pre-seed · $180k

See the model run on your side of the wall.

A 30-minute call, a live node, and the log. Everything on this page is reproducible from the repository.