Skip to content

🔥 URGENT: GF16 integration into train_gpt_mlx.py — NOW #1

Description

@gHashTag

Задача — выполнить СЕЙЧАС

Шаг 1: Проверить структуру gf16.py

cat submissions/v03_gf16/gf16.py | head -50

Найти класс GF16Layer или функции compress_weights / phi_encode.


Шаг 2: Создать train_gpt_mlx_gf16.py

Скопировать train_gpt_mlx.py → submissions/v03_gf16/train_gpt_mlx_gf16.py

Добавить GF16 quantization после каждого optimizer step:

# В конец train loop, после optimizer.step():
import sys
sys.path.insert(0, 'submissions/v03_gf16')
from gf16 import compress_weights, phi_encode

# После каждого optimizer step — квантовать веса
def apply_gf16_quantization(model):
    """Apply GF16 phi-based quantization to all linear layers."""
    PHI = 1.6180339887498948482
    for name, param in model.parameters():
        if hasattr(param, 'shape') and len(param.shape) == 2:
            # Scale weights by phi before quantization
            scaled = param * PHI
            # Quantize to 8-bit grid aligned to phi
            quantized = mx.round(scaled * 127) / 127 / PHI
            param[...] = quantized

Шаг 3: Добавить --gf16 флаг

import argparse
parser = argparse.ArgumentParser()
parser.add_argument('--gf16', action='store_true', help='Enable GF16 quantization')
parser.add_argument('--steps', type=int, default=200)
args = parser.parse_args()

Шаг 4: Запустить сравнение

cd ~/parameter-golf-trinity
source .venv/bin/activate

# Baseline (уже запущен или запустить заново):
python train_gpt_mlx.py --steps 200 2>&1 | tee experiments/baseline_200.log

# GF16 версия:
python submissions/v03_gf16/train_gpt_mlx_gf16.py --steps 200 --gf16 2>&1 | tee experiments/gf16_200.log

# Сравнить финальный loss:
tail -5 experiments/baseline_200.log
tail -5 experiments/gf16_200.log

Шаг 5: Commit результаты

git add submissions/v03_gf16/train_gpt_mlx_gf16.py
git add experiments/baseline_200.log
git add experiments/gf16_200.log
git commit -m "exp: GF16 vs baseline MLX comparison — step 200"
git push origin main

Ожидаемый результат

Метрика Baseline GF16
Final loss @200 ~? ~?
Model size 100% ~50%
BPB estimate ~1.22 target <1.20

Definition of Done

  • train_gpt_mlx_gf16.py создан с --gf16 флагом
  • Оба лога записаны в experiments/
  • Финальный loss обоих run задокументирован
  • Commit запушен в main

Deadline: сегодня ночью.

Activity

  1. gHashTag commented on Jun 14, 2026

    @gHashTag
    OwnerAuthor

    Status update (2026-06-15) -- honest DoD reconciliation + a new planning artifact

    Negative-first. Revisiting this issue's Definition of Done against what is actually in the repo today (HEAD pushed 2026-04-30):

    • submissions/v03_gf16/train_gpt_mlx_gf16.py exists (it is in the tree, ~50 KB).
    • --gf16 argparse flag: NOT confirmed in the committed file (a def main() is present but I do not see the add_argument('--gf16', ...) from Step 3). Needs a check or a follow-up commit.
    • experiments/baseline_200.log and experiments/gf16_200.log: NOT in the repo. Only zero-byte smoke_001.log / smoke_mlx_001.log are committed, so the GF16-vs-baseline final-loss comparison table is still empty (~? / ~?).
    • Final loss of both runs documented: NOT done (depends on the two missing logs).

    So the DoD is PARTIALLY complete: the trainer file landed, but the measured comparison that was the point of the task did not. The README's "Status: GF16 prototype working / MLX smoke test in progress" is consistent with that -- prototype yes, measured result no.

    New artifact (planning, not a result). For deciding whether a sub-1M run is even worth the compute, I published a small deterministic Rust calculator, floor_oracle, as a public gist:
    https://gist.github.com/gHashTag/5329efee1f6af8fb8e9694546600158b (Apache-2.0, self-test 20/20).

    It does not measure a model and it makes no capability claim (its fields stay TRAIN_BOX_PENDING). It computes, from a config's own constants: exact param count (reproduces a 493,056-param anchor two ways), the Chinchilla deficit at 20 tok/param, the multiplicative gap to published floors on BOTH params and tokens (e.g. TinyCodeLM-150M ~304x / ~1e6x, SmolLM-135M ~274x), and the Wilson 95% 0-success eval-power ceiling (n=4 -> ~0.49, n=16 -> ~0.19, n=32 -> ~0.11). The floor numbers are typed from the papers and carry their arXiv ids; the self-test checks the math and determinism, not the literature.

    Note on the SOTA anchor: this repo's README cites 1.0810 BPB (bigbag) as current SOTA; I should reconcile the gist/notes against that number rather than an older ~1.06 figure -- flagging so the two do not drift.

    Next concrete step to actually close this issue: either confirm/add the --gf16 flag and commit the two *_200.log files with the final-loss table filled in, or relabel the issue as superseded if the MLX track is parked. No hype, no deadline pressure -- just an honest reconciliation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions