Repository navigation
🔥 URGENT: GF16 integration into train_gpt_mlx.py — NOW #1
Description
Activity
Status update (2026-06-15) -- honest DoD reconciliation + a new planning artifact
Negative-first. Revisiting this issue's Definition of Done against what is actually in the repo today (HEAD pushed 2026-04-30):
-
submissions/v03_gf16/train_gpt_mlx_gf16.pyexists (it is in the tree, ~50 KB). -
--gf16argparse flag: NOT confirmed in the committed file (adef main()is present but I do not see theadd_argument('--gf16', ...)from Step 3). Needs a check or a follow-up commit. -
experiments/baseline_200.logandexperiments/gf16_200.log: NOT in the repo. Only zero-bytesmoke_001.log/smoke_mlx_001.logare committed, so the GF16-vs-baseline final-loss comparison table is still empty (~?/~?). - Final loss of both runs documented: NOT done (depends on the two missing logs).
So the DoD is PARTIALLY complete: the trainer file landed, but the measured comparison that was the point of the task did not. The README's "Status: GF16 prototype working / MLX smoke test in progress" is consistent with that -- prototype yes, measured result no.
New artifact (planning, not a result). For deciding whether a sub-1M run is even worth the compute, I published a small deterministic Rust calculator,
floor_oracle, as a public gist:
https://gist.github.com/gHashTag/5329efee1f6af8fb8e9694546600158b (Apache-2.0, self-test 20/20).It does not measure a model and it makes no capability claim (its fields stay TRAIN_BOX_PENDING). It computes, from a config's own constants: exact param count (reproduces a 493,056-param anchor two ways), the Chinchilla deficit at 20 tok/param, the multiplicative gap to published floors on BOTH params and tokens (e.g. TinyCodeLM-150M ~304x / ~1e6x, SmolLM-135M ~274x), and the Wilson 95% 0-success eval-power ceiling (n=4 -> ~0.49, n=16 -> ~0.19, n=32 -> ~0.11). The floor numbers are typed from the papers and carry their arXiv ids; the self-test checks the math and determinism, not the literature.
Note on the SOTA anchor: this repo's README cites 1.0810 BPB (bigbag) as current SOTA; I should reconcile the gist/notes against that number rather than an older ~1.06 figure -- flagging so the two do not drift.
Next concrete step to actually close this issue: either confirm/add the
--gf16flag and commit the two*_200.logfiles with the final-loss table filled in, or relabel the issue as superseded if the MLX track is parked. No hype, no deadline pressure -- just an honest reconciliation.-
Задача — выполнить СЕЙЧАС
Шаг 1: Проверить структуру gf16.py
cat submissions/v03_gf16/gf16.py | head -50Найти класс
GF16Layerили функцииcompress_weights/phi_encode.Шаг 2: Создать train_gpt_mlx_gf16.py
Скопировать
train_gpt_mlx.py→submissions/v03_gf16/train_gpt_mlx_gf16.pyДобавить GF16 quantization после каждого optimizer step:
Шаг 3: Добавить --gf16 флаг
Шаг 4: Запустить сравнение
Шаг 5: Commit результаты
git add submissions/v03_gf16/train_gpt_mlx_gf16.py git add experiments/baseline_200.log git add experiments/gf16_200.log git commit -m "exp: GF16 vs baseline MLX comparison — step 200" git push origin mainОжидаемый результат
Definition of Done
train_gpt_mlx_gf16.pyсоздан с --gf16 флагомexperiments/Deadline: сегодня ночью.