Physix Frontier · News Briefing Card (enbrief · Oct 2, 2026)

He Kaiming's team solves ARC with ImageNet-pretrained vision models

KEY FACTS

  • He Kaiming's team proposed NAT-ARC, a purely vision-based approach to ARC abstract reasoning tasks that does not rely on large language models.
  • The method first uses an ImageNet MAE-pretrained encoder, then trains on the ARC training set, and applies LoRA fine-tuning per task at test time.
  • The single NAT-ARC model scores 63.4% pass@2 on ARC-1, reaching 70.2% with an ensemble.
  • Against the specialized LLM approach The ARChitects at 71.6%, NAT-ARC uses roughly one quarter of its parameter count.
  • The paper found that after pretraining the model's scaling curve turns positive, whereas without pretraining larger models previously performed worse.

KEY DATA

63.4%Single-model ARC-1 pass@2
70.2%Ensemble ARC-1 pass@2
0.6BSingle-model parameter count
about 2BTotal ensemble parameter count

PHYSIX OBSERVATION

The pure-vision route uses pretraining to patch up its scaling weakness, and with far fewer parameters than LLMs it approaches specialized systems, showing that abstract reasoning does not necessarily depend on language. For the industry, this suggests visual priors can transfer to symbolic tasks, that ARC solutions may no longer be monopolized by LLMs, and that small models plus pretraining deserve fresh evaluation.

Source: enbrief original report ↗