This project studies whether a recurrent block in Qwen/Qwen2.5-0.5B-Instruct can learn one pointer transition per loop, generalize to deeper tasks, and eventually repair mistakes while preserving correct answers.
The current focus is further pointer training and depth generalization. Full-loop tests show strong execution on unseen mappings at trained depths 1–6, with limited extension beyond them. Overscaling tools are prepared for later; those experiments have not run. See current status and the research plan.
Use Python 3.11 and run commands from the repository root. Windows/NVIDIA users should follow the WSL2/CUDA setup guide first; other environment details are in setup.
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python -m pip checkDownload the base revision used for the reproducible experiments:
hf download Qwen/Qwen2.5-0.5B-Instruct --revision 7ae557604adf67be50417f59c2c2f167def9a775Check ordinary model loading and generation:
python scripts/smoke_test_qwen.pyActivate .venv in each new terminal. The commands below default to CPU unless a device is supplied. Use --device mps on Apple Silicon or --device cuda on the configured NVIDIA desktop. For deterministic CUDA runs, set:
export CUBLAS_WORKSPACE_CONFIG=:4096:8Dataset code lives in scripts/dataset/. Preview five examples without writing files:
python -m scripts.dataset --seed 17 --dry-runUse --dry-run 10 for ten examples. Reproduce the existing dataset with master seed 17, the pinned tokenizer above, and this complete configuration:
python -m scripts.dataset \
--seed 17 \
--train-count 10000 \
--validation-count 1000 \
--test-count 1000 \
--depth-test-count 1000 \
--min-depth 1 \
--max-train-depth 8 \
--max-eval-depth 16 \
--output data/pointer/seed-17
python -m scripts.dataset --verify data/pointer/seed-17The existing seed-17 dataset already reaches depth 16; no regeneration is needed for depth-6 training. Generated files are excluded from Git, so use the reproduction command above only if the dataset is missing on a new machine.
File under data/pointer/seed-17/ |
Examples | Depths | Examples per depth |
|---|---|---|---|
train.jsonl |
10,000 | 1–8 | 1,250 |
validation.jsonl |
1,000 | 1–8 | 125 |
test.jsonl |
1,000 | 1–8 | 125 |
depth_test.jsonl |
1,000 | 9–16 | 125 |
The depth-6 training config selects only depths 1–6 from train.jsonl. Its validation_max_depth: 8 controls training-time monitoring; the paired depth evaluator separately reads depth_test.jsonl and runs through depth 16. See dataset details and paired evaluation.
Inspect three ordinary-model questions before running the full test:
python -m scripts.eval.naive_test --model Qwen/Qwen2.5-0.5B-Instruct --test
python -m scripts.eval.naive_test --model Qwen/Qwen2.5-0.5B-InstructThe ordinary baseline uses three-shot examples at depths 1, 2, and 3, defined in prompts/pointer_task.txt. --model also accepts a saved model directory. Models load locally by default; --download permits missing downloads.
Progress and throughput appear in the terminal. Predictions and summaries go to eval/pointer_task/. See baseline evaluation for loading and scoring options.
Preview the initial training configuration before launching it:
python -m scripts.training.train_pointer --config configs/stage1_pointer.json --dry-run
python -m scripts.training.train_pointer --config configs/stage1_pointer.json --device cudaConfigs live in configs/. The trainer uses raw dataset prompts and exact per-loop supervision. A compact Rich dashboard shows progress, ETA, losses, accuracy, and memory. Checkpoints and logs go to models/stage1_pointer/; model binaries are excluded from Git.
The next experiment uses configs/stage1_pointer_depth6_loopbalanced.json: fresh adapters on the existing 30,000 mappings, with equal loss weight per loop position. Follow the loop-balanced training commands, starting with the dry-run. The training guide also covers checkpoint selection, resume, tmux, and old-artifact cleanup.
Pass a complete saved step directory, including its adapter weights and tokenizer:
python -m scripts.eval.loop_test \
--model models/stage1_pointer/<run>/step-000625 \
--device cuda --loops 8 --testRemove --test for all 1,000 test examples. Outputs go to eval/pointer_loops/. This reads the model after every recurrent loop; the ordinary three-shot prompt is not used. See full-loop evaluation for commands and metrics.
Preview the separate absorbing-terminal task variant without loading a model or writing files:
python -m scripts.eval.overscaling_test --dry-runCheckpoint sweeps, repair/damage scoring, and survival exports are available but running them is deferred until further Stage 1 training and review. See overscaling usage.