Compression
How small the theta bundle is versus the original projection weights, measured across LoRA ranks at real Gemma 3 1B dimensions.
Compression is the reason weightless inference exists, so it’s the first thing to measure. The metric is the size of the theta bundle (W0 + s + U + V across every layer and family) as a fraction of the original projection weights it replaces. The rank of the LoRA delta sets both the fidelity and the size, so the ratio is always read against a rank.
Bundle size versus rank
The theta bundle stays well under the original at every rank tested. It grows with rank because only U and V scale with it; W0 and s are fixed.
The full sweep, with the original projection weights at 1778.91 MB (float32):
| Rank | Theta bundle | Ratio of original |
|---|---|---|
| 1 | 72.05 MB | 4.05% |
| 4 | 79.20 MB | 4.45% |
| 8 | 88.73 MB | 4.99% |
| 16 | 107.80 MB | 6.06% |
| 32 | 145.92 MB | 8.20% |
| 64 | 222.17 MB | 12.49% |
At the rank 4 serving default, the bundle is 4.45% of the original projection weights. Even at rank 64 it stays under an eighth.
Where the bytes go
Breaking the rank 32 bundle into its four components shows why the ratio stays low: the shared block W0 is stored once per family, the scale s is negligible, and only the LoRA factors U and V carry per-layer cost.
| Component | Size | Share of bundle |
|---|---|---|
W0 (shared blocks) |
68.42 MB | 46.9% |
s (per-layer scales) |
1.25 MB | 0.9% |
U (LoRA left factors) |
40.04 MB | 27.4% |
V (LoRA right factors) |
36.21 MB | 24.8% |
W0 is nearly half the bundle at rank 32 and is fixed regardless of rank, which is why dropping to rank 4 barely moves the total: the LoRA factors shrink but the shared block doesn’t. Raising G (more shared blocks) adds W0 copies, trading bundle size for a closer per-band fit.