Revision history of "Using Bonsai QWEN Models"

Jump to navigation Jump to search

Diff selection: Mark the radio buttons of the revisions to compare and hit enter or the button at the bottom.
Legend: (cur) = difference with latest revision, (prev) = difference with preceding revision, m = minor edit.

  • curprev 09:19, 27 August 2026PeterHarding talk contribs 2,284 bytes +23
  • curprev 09:18, 27 August 2026PeterHarding talk contribs 2,261 bytes +2,261 Created page with " See - https://www.youtube.com/watch?v=V6LmF7TuBmY (Bonsai 27B Runs Qwen 3.6 27B at 10x less memory) Not a corrupt download — a format mismatch. Q2_0 in that filename isn't stock llama.cpp's quant; it's Q2_0_g128, PrismML's custom ternary packing (2-bit slots, one FP16 scale per 128 weights) that only their fork's kernels understand. Your stock build computes a different byte size for the ternary tensors, so its running offset drifts from what's written in the header,..."