Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
Ranking
Overall
82
Content
95
Popularity
52
Observed public metrics from 1 member.
Merged summary
TL;DR - Logic pre-pretraining on formal derivations helps language models acquire linguistic skills faster while producing representations that are easier to compress. At 100B-token scale, it reaches 80% linguistic-task accuracy with 36B fewer tokens than standard initialization.
- Formal derivations introduce structural biases involving variables, quantifiers, relational dependencies, and predicate-argument composition.
- Logic-PPT outperforms Dyck-language and procedural-algorithm pre-pretraining baselines.
- It produces persistent, lower-rank, spectrally concentrated representations.
- Pruned models match dense-baseline performance at approximately 33% sparsity.
Sources (1)
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
Public signals
Hugging Face upvotes 10
TL;DR - Logic pre-pretraining on formal derivations helps language models acquire linguistic skills faster while producing representations that are easier to compress. At 100B-token scale, it reaches 80% linguistic-task accuracy with 36B fewer tokens than standard initialization.
- Formal derivations introduce structural biases involving variables, quantifiers, relational dependencies, and predicate-argument composition.
- Logic-PPT outperforms Dyck-language and procedural-algorithm pre-pretraining baselines.
- It produces persistent, lower-rank, spectrally concentrated representations.
- Pruned models match dense-baseline performance at approximately 33% sparsity.