🛰️ Daily AI Frontier
‹ back to 2026-08-05

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

arXiv cs.CL LLMs & Foundation Models Jo-Ku Cheng, Nikolaos Aletras, Marco Valentino 2026-08-04

TL;DR - Logic pre-pretraining on formal derivations helps language models acquire linguistic skills faster while producing representations that are easier to compress. At 100B-token scale, it reaches 80% linguistic-task accuracy with 36B fewer tokens than standard initialization.

  • Formal derivations introduce structural biases involving variables, quantifiers, relational dependencies, and predicate-argument composition.
  • Logic-PPT outperforms Dyck-language and procedural-algorithm pre-pretraining baselines.
  • It produces persistent, lower-rank, spectrally concentrated representations.
  • Pruned models match dense-baseline performance at approximately 33% sparsity.

view merged work →