NORTH-STAR · 북극성
Two objects, one system: ① a ruler that defines what "correct" means, and ② GRAM, a weight-shared recursion that builds the answer. The old thesis kept ① hand-written (Ec). The current thesis learns ① — order-MDL (permute the solving order M ways; one shared model can't store M accidents, so only the rule survives) × parametric-MDL (a per-puzzle residual latent kept short by a code-length prior, weights frozen) — from few examples, no augmentation. The hand rule is demoted to oracle baseline and sound verifier.
① Ruler · Rφ LEARNED
order-MDL × parametric-MDL
A GRAM instance trained to predict order-invariant legal-sets across permuted solving orders; eps-KL is the description-length knob. kstep legal-F1 0.993, most order-invariant arm.
② GRAM · recursion & search
the consumer
Weight-shared K-step recursion + verified best-first search. Learns from dense supervision; acceptance is always sound verification (wrong-accepts 0 everywhere).
The consumption-locus law. A learned signal lives or dies by where it is consumed, not what it contains. Gradient descent on energy died 8×. The hard-candidate seat dies even at F1 0.99 (one pruned true digit kills the branch — the seat demands ~100%). Seats that average the signal live: observation ✓, branch ordering ✓ (+8~12pp, no training), dense teaching ← Stage G, next.
IRED — learned energy, descended
take: energy can be learned from data.
differ: we never descend it. Descent died 8× here (two-basin, non-navigable). Energy/ruler is consumed as observation, ordering, and target — discrete seats only.
CompressARC — MDL as the whole method
take: description length is the right few-data objective.
differ: MDL is not our solver, it is the trainer of the ruler — a two-part code: shared rule weights (order-MDL) + a short per-puzzle residual (parametric-MDL).
Test-time training (ARC-style TTT)
take: adapt at inference, per task.
differ: we adapt only a throwaway latent zH under a code-length prior — weights frozen, no committed update, discarded after the puzzle. First run: weakOOD solve +5pp.
Classical CSP — MRV, propagation, verification
take: sound machinery is non-negotiable; acceptance is always verified.
differ: the hand rule is the oracle baseline, not the endgame — learned parts take the seats a hand rule can't fill on ARC/crystal, where no hand rule exists.
Energy as a gradient field — descend −∇E to the answer.
Verifier perfect (AUC 1.0), navigation impossible: spurious low-energy basins trap every chain; fine-tuning the field breaks ranking instead (flow ↔ ranking coupled). Died 8× across levers. Lesson: a verifier is not a navigator.
Per-step deep supervision — decode→denoise→regress at every step.
Endpoint-only training 0.24, dense per-step 0.41 OOD. Lesson: dense supervision wins where endpoints fail — the seed of every later win.
Gold-free teaching works — hand-energy as the only teacher.
GRAM trained without gold answers generalizes with zero train→held-out gap, held-out 0.47 > gold-taught 0.296. Lesson: the teaching pipeline exists; what's missing is a learnable teacher.
Energy as observation — first net-positive in the recursion.
The violation-map fed as a detached input channel (never a gradient): train exact 7×, OOD ceiling broken. Same info as a 13-ch crutch lowers the ceiling. Lesson: value is set by the consumption seat.
CA-1 operator learning — first held-out liftoff.
Initial-state diversification + per-step re-damage breaks the exposure-bias wall: held-out exact 0 → 0.09. Lesson: cover the operator's input domain, not more epochs.
①×② verified search — 0.31 held-out / 0.75 weakOOD.
Model orders MRV branches (worth ~4× budget), hand rule prunes, sound verification accepts; wrong-accepts 0. Distilling traces back overfits (weakOOD −9pp). Lesson: search-time consumption beats weight-time distillation off-distribution.
Learned scalar energy as search value / reward / latent drift — all die.
S1 ranking-champion (AUC 0.988) fails to convert to solving; S2 graded rewards collapse CA learning; S4 latent drift worsens every axis. Meanwhile no-training value-guided search wins +8~12pp near-distribution. Lesson: scalar-energy learning hit its ceiling on Sudoku; the learnable signal that remains is structured (per-cell), not scalar.
Order-MDL ruler built; the easy target saturates.
Onestep (local exclusion) — every arm F1 ≈ 1.0, no discrimination; KL even over-compresses (1.0→0.30). Promoted to kstep (propagation fixpoint), where headroom exists. Lesson: a testbed must have headroom before it can falsify.
The ruler is learned — C2 (orders×4 + eps+KL) F1 0.993, most order-invariant.
Deterministic arms flatline 6240 steps at the saddle; eps noise + MDL compression escapes and wins. Sign flip: the same KL that hurt the easy target carries the hard one. Lesson: noise + compression is the key to hard-target recursion training.
…but it cannot sit in the hand rule's seat: coupling solve ≈ 0 vs 0.30/0.66.
Cell-level probe: off-path hypothesis half-refuted (F1 holds off-path); the killers are 12% true-digit FN and 18% false dead-ends on clean states — (1−0.125)50 ≈ 0.1% branch survival. Lesson: hard-decision seats demand ~100%; the learned ruler's value lives where no hand rule exists — teach, don't decide. → Stage G.
ARC 무증강 생성은 시연을 0.9877로 외우고 새 입력은 0.012 — 재귀 깊이는 원인이 아니다.
K=1/2/4/6 전부 test 내용 0.644~0.671로 부동. 헌법이 처방한 암기-견제 3스코프(파라미터·손상·잠재 MDL) 중 ARC에 적용된 것 0건.
5팔 사전등록: 학습 0회의 고전 단일-호 일관성 룩어헤드(SAC-lite) 0.1353 > 학습 게이트 0.0677.
주 추정량 −0.0677, one-sided 95% 하한 −0.0902. 07-19의 end-to-end 확증 철회. Lesson: 학습 성분 주장에는 그 성분이 흉내 내는 고전 팔이 통제군으로 반드시 들어가야 한다.
크리스탈 2×2 사전등록: 잠재 압축 −0.0270, 손상복구 −0.0090 — 헌법의 중심 가설이 첫 직접 시험에서 기각.
"압축이 증강의 대체물"이라는 명제의 첫 요인 분해 시험. 두 주효과 모두 음수.
생성기 단독 순기여 0 (고전 전파의 진부분집합) + 선언된 눈금과 채택된 눈금이 다른 물건이었다.
P-SOLO: 생성기만 푼 퍼즐 0/2000 · 0/300 · 0/400. 같은 체크포인트에서 선언된 판정식은 +0.000(정의 불능), 채택된 판정식은 +0.785. Lesson: 칸별 정확도는 해결률의 대리 변수가 아니다 — 다음 표적은 "언제 커밋할지". → 2026-07-26 페이지.
칸별 정확도 0.411로는 빈칸 약 56개짜리 한 판을 풀 수 없다 — 50%에 필요한 값은 약 0.988이다.
사전등록 P-SOLO에서 생성기 단독 해결률은 .0220 / .0000 / .0150이고, 생성기가 푼 집합은 고전 전파가 푸는 집합의 진부분집합이다(생성기만 푼 것 3표본 전부 0). 같은 체크포인트에서 발사 스크립트가 선언한 판정식은 +0.000(정의 불능), 원장이 채택한 판정식은 +0.785로 갈린다. 다음 표적은 정확도가 아니라 커밋 결정이다.
국소 확신도는 상위 구간에서 신뢰할 수 없는 선택 신호였다 — 그리고 우리 최고 기록 0.865는 MRV 대조군과 동률이었다.
자기부정 통제: 신경 분기-순서를 고전 휴리스틱으로 바꿔도 0.865, 예산별 차이 다섯 개 모두 CI가 0을 포함. 전역 일관성 검사를 결합하면 커밋 정밀도가 0.995까지 오르지만 그 대부분은 전파 특징 단독(0.982)으로 설명된다. end-to-end 개선은 검색을 끈 조건에서만 확인됐고, 크리스탈 이식 주장(AUC 0.315)은 fold 아티팩트로 철회. 2026-07-20 5팔 통제로 §5의 end-to-end 확증은 철회됐다 — 페이지 상단 정정 배너 참조.
The ruler is now learned: order-MDL trains Rφ to F1 0.993 — and the honest negative that redirects the program to teaching.
kstep ladder: deterministic arms flatline, eps+KL escapes the saddle and wins with the best order-invariance. Hard-candidate coupling fails at F1 0.99 (12% true-digit FN diagnosed at the cell level) → the consumption-locus law extended; parametric-MDL TTT first run +5pp weakOOD; Stage G (ruler-as-teacher, recovery ratio) is the next verdict.
Operator learning lifts held-out exact 0 → 0.09; ①×② verified search takes it to 0.31 / 0.75.
Init-diversification breaks the exposure-bias wall; energy stays a detached observation; the model orders MRV branches, Ec prunes, sound verification accepts — project record, wall opened.
The dead "descend the energy" thesis.
Pathwise −∇E into the generator: verifier AUC 1.0, gold E < random, train exact ≈1.0 — falsified as the dead gradient path. Kept as honest history, not as a live method.
정직한 상태. 학습된 자는 실재한다(k스텝 legal-F1 0.993, 순서 불변). 그러나 2026-07-20 이후 사전등록 통제 아래에서 스도쿠·크리스탈·ARC 세 도메인 모두 학습 성분의 최종 지표 순기여가 0 또는 음수로 측정됐고, 2026-07-26에는 생성기 단독 해결률의 순기여도 0으로 확정됐다(고전 전파의 진부분집합). 헤드라인 눈금을 칸별 회복 비율에서 해결률 + 커밋 정밀도@커버리지로 옮기는 개정이 사용자 승인 대기 중이다. 남은 미검정 축은 마지막 사이클 안쪽 스텝에 손실을 거는 자리 하나이며, 그마저도 "타깃이 스텝마다 다를 때"만 검정 가치가 있다.