🛰️ Daily AI Frontier
‹ back to 2026-08-29

SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models

Research LLMs & Foundation Models

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models

Merged summary

TL;DR - SCIT is a causal testing protocol for identifying which transformer cache components carry computations in latent chain-of-thought models. It finds that arithmetic reasoning in tested GPT-2 checkpoints primarily transfers through value-cache suffix trajectories, while larger or non-arithmetic models can use different cache regions.

  • SCIT patches exact cache segments between source and recipient examples, combining sufficiency, necessity, component-split, corruption, and control tests.
  • CODI-GPT2 showed sufficiency-and-necessity evidence that late value-cache suffixes carry counterfactual arithmetic computations.
  • A Sim-CoT-style GPT-2 reproduced the sufficiency pattern, but lacked enough matched-corruption evidence to establish necessity.
  • Carrier mechanisms varied with model competence and task: some cells used latent-tail values/KV, while competent 8B and repaired non-arithmetic cells relied on prompt-prefix or full-cache K/V.

Sources (1)

SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models

arXiv cs.CL Yi Ding, Lijun Huang, Menglin Yang 2026-08-27 arXiv:2608.27265
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:17:49.573549 UTC

TL;DR - SCIT is a causal testing protocol for identifying which transformer cache components carry computations in latent chain-of-thought models. It finds that arithmetic reasoning in tested GPT-2 checkpoints primarily transfers through value-cache suffix trajectories, while larger or non-arithmetic models can use different cache regions.

  • SCIT patches exact cache segments between source and recipient examples, combining sufficiency, necessity, component-split, corruption, and control tests.
  • CODI-GPT2 showed sufficiency-and-necessity evidence that late value-cache suffixes carry counterfactual arithmetic computations.
  • A Sim-CoT-style GPT-2 reproduced the sufficiency pattern, but lacked enough matched-corruption evidence to establish necessity.
  • Carrier mechanisms varied with model competence and task: some cells used latent-tail values/KV, while competent 8B and repaired non-arithmetic cells relied on prompt-prefix or full-cache K/V.
item →