Verified low-bit inference for the CPU you already have — the documentation, as of 0.2.7.
gate record 0.3.0: Savante answered by bankML’s own forward pass — whole conversations identical to llama-server on 9 / 9 turns (text, token counts, prompt-cache reuse) · own forward pass 1,064 / 1,064 rows bit-exact for the 1-bit and the ternary model · seeded sampling identical on 40 / 40 continuations · 8,188,239,872 ternary weights and 762 / 762 dot products bit-exact against llama.cpp b11192’s compiled library · one ternary token 0.222 s vs 2.139 s