HyperAIHyperAI

Command Palette

Search for a command to run...

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin Ruixuan Deng Junran Wang

Abstract

Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4–3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen38B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https: //lunamos.github.io/stop-thinking-too-early/.

One-sentence Summary

The authors show that inserting a task-trained rank-8 LoRA at one early layer while keeping all model weights frozen corrects pretrained transformers' tendency to stop following chain references too early: Qwen38B improves from 15.5% to 99% exact accuracy on 24-line chains, Ouro-1.4B reaches 60 to 160 lines, and task-specific LoRAs also improve MuSiQue.

Key Contributions

  • The paper quantifies the default reference-following computation in pretrained transformers, showing across thirteen base models and two looped families that models reliably follow only 1.4 to 3.6 lines and that extra pretrained loops add little.
  • The paper introduces a lightweight intervention that keeps model weights frozen and adds a task-trained rank-8 LoRA at one early layer, improving Qwen38B from 15.5% to 99% exact accuracy on 24-line chains, reaching 50 lines with longer training, and allowing Ouro-1.4B to follow 60 lines after four loops and at least 160 after eight.
  • The paper provides a causal explanation of the extended computation as a relay in which chain identity is passed through a short range of middle layers, and parent-line attention is required for frozen heads to read further up the chain. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models, and early LoRAs improve MuSiQue exact match.

Introduction

Language models are often expected to follow in-context reference chains, such as resolving a sequence of variable assignments before answering a query. This ability matters for multi-hop reasoning and instruction following because the model must compose facts supplied entirely in context. The authors find that frozen pretrained models reliably track only about 1.4 to 3.6 assignment lines, and even large models or pretrained looped architectures do not extend this reach much, although transformers can in principle perform much longer reference-following computations. Their main contribution is showing that this failure comes from a short default computation in middle layers, and that a very small adaptation, such as a rank-8 LoRA at one layer, can extend the relay to 50 lines in one pass and over 160 lines with loops. They also provide a causal account of the relay and a prospective placement test for where such adaptations should be applied.

Dataset

The authors use a synthetic dataset of reference-chain programs. Each program contains ccc chains, each with ddd assignments, followed by a query and Output:.

  • Sources and composition: The data are generated rather than collected from an existing corpus. Names, nouns, and queried chains are randomized. Each root assignment stores a single-token noun, and every later assignment names the preceding variable. Each assignment occupies a separate prompt line. The answer is the queried chain’s root noun.
  • Processing and orderings: Two orderings are used: level order, where assignments are grouped by depth and shuffled within each level, and interleaved order, where chains are randomly merged while preserving definition before use. A line’s pointer is the variable on its right-hand side, and its parent line defines that variable. The prompt header, vocabulary, training details, and evaluation protocols are given in Appendix A.
  • Scale and schema: Chain length counts assignments including the root. The example mixes two chains of three lines, while the general setup uses ccc chains and ddd assignments. Headline standard results use 200 programs per cell; other sample sizes are specified in the appendix.
  • Evaluation metrics: Choice accuracy selects among the chains’ root values, with chance 1/c1/c1/c. Exact accuracy requires the correct root to rank first over the full vocabulary. Reach is the longest chain followed with at least 80% accuracy, linearly interpolated at the first downward crossing.
  • How the data is used: Standard models are evaluated with three-chain choice accuracy. Standard-model LoRA evaluations use two-chain exact accuracy. Looped-model experiments use two-chain choice accuracy unless stated otherwise. The task is also used for LoRA training, where training combines answer cross-entropy with a KL penalty on WikiText-103. Standard Qwen LoRA trains through 20 lines; Ouro trains with four loops; longer-trained versions see up to 40 lines and a larger vocabulary; Huginn sees up to 24 lines.

Method

The authors design a diagnostic reference-chain task and pair it with a minimal, single-layer intervention in a pretrained model. The goal is to isolate whether long-range copying and variable binding can be recovered by editing one hidden-state transformation while keeping the rest of the model frozen.

A program contains ccc chains, each with ddd assignments. The root assignment stores a single-token noun, and every later assignment in a chain names the preceding variable. Chain length counts the root plus all subsequent assignments. For example, a chain may start with K = apple, continue with B = K, and then D = B, so the correct printed value is apple. The final query asks the model to print one of the variables, and the answer must be selected from the root values of the chains. Names, nouns, and queried chains are randomized. Programs can be arranged in level order, where assignments are grouped by depth and shuffled within each level, or in interleaved order, where chains are merged while preserving the constraint that a variable is defined before use. Each right-hand side variable acts as a pointer to the line where that variable was defined.

The main intervention is a rank-8 LoRA applied to the residual stream at the input of a single layer. The transformation is

h←M(h)=s h+BAh,h \leftarrow M(h) = s\,h + B A h,h←M(h)=sh+BAh,

where hhh is the hidden state at the selected layer input, A∈R8×nA \in \mathbb{R}^{8 \times n}A∈R8×n, B∈Rn×8B \in \mathbb{R}^{n \times 8}B∈Rn×8, and sss is a learned scalar with no bias term. This is a representation intervention in the DiReFT form. It acts independently at each token position, while the frozen attention and MLP layers continue to handle all token-to-token communication. Only AAA, BBB, and sss are trained.

The learned scale remains close to identity. Standard Qwen and Ouro LoRAs learn s=1.0006s = 1.0006s=1.0006 and s=1.004s = 1.004s=1.004, respectively, and their longer-trained variants learn 1.0081.0081.008 and 1.0021.0021.002. Thus the intervention remains a small rank-8 perturbation of the original state. The number of trainable parameters is 65,53765{,}53765,537 for Qwen3-8B, 32,76932{,}76932,769 for Ouro-1.4B, and 84,48184{,}48184,481 for Huginn. For looped models, the same LoRA is applied in every loop unless stated otherwise.

Training combines answer cross-entropy with a KL penalty, KL(p0∥pM)\mathrm{KL}(p_0 \parallel p_M)KL(p0​∥pM​), computed on WikiText-103. This regularizer keeps the modified model close to the original model on general text. In the standard setting, the Qwen LoRA enters at layer 14 and trains on programs up to 20 lines, while the Ouro LoRA enters at layer 6 and trains with four loops. Longer-trained Qwen and Ouro LoRAs see up to 40 lines and a larger vocabulary, and the Huginn LoRA sees up to 24 lines. The WikiText perplexity changes only slightly, from 10.14 to 10.148 for Qwen and from 13.294 to 13.302 for Ouro at four loops, indicating that the method is a controlled representation-level edit rather than a broad behavior change.

Experiment

The experiments evaluate both frozen language models and LoRA interventions on synthetic pointer-chain programs plus multi-hop question answering, using causal tracing, attention analysis, and placement sweeps. Default models resolve only a few reference steps before the relay stops, regardless of added depth or recurrence. A small LoRA placed in specific middle layers enables much longer chains by restarting and extending an existing relay mechanism, and the same placement preference transfers to improved multi-hop question answering. The findings show that useful computation is accessible under targeted interventions, but its success depends sharply on where and when the intervention is applied.

Across several adaptation methods on two-chain, 24-line programs, accuracy is substantially higher when the edit is placed at layer 14 than at layer 26. The pattern holds for a small LoRA, a larger projection LoRA, and FLAS-style flow and block edits with matched training and step counts. The early placement advantage is therefore shared across edit types and is not simply a function of parameter count. All tested interventions show a large advantage at the earlier layer, with later-layer accuracy dropping to a much lower band. The smallest LoRA achieves the best early-layer result despite having far fewer parameters than the projection and block alternatives, underscoring that placement matters more than edit capacity here.

The frozen cutoff predicts where LoRA updates remain effective, matching the observed placement bracket in three of four held-out models and missing one held-out model by five layers from the bracket midpoint. Its mean error is lower than the 45% depth and value-copy baselines and essentially tied with a post-hoc relative depth rule. Placement itself is sharp, as moving a Qwen3-8B LoRA one layer later can reduce reach from a large extension to near the frozen baseline. The frozen cutoff correctly locates the working placement bracket in three of four held-out models, with mean absolute error lower than 45% of depth and value-copy baselines. A one-layer shift from layer 20 to 21 in Qwen3-8B sharply reduces reach, and later-layer-only interventions show remaining depth alone does not determine whether useful computation can follow.

The experiments evaluate where edits and LoRA updates should be placed, using two-chain 24-line programs and held-out models. Across small LoRA, projection LoRA, and FLAS-style edits, placing updates at an earlier layer consistently yields substantially higher accuracy than at a later layer, with the smallest LoRA performing best early despite fewer parameters, indicating that placement matters more than edit capacity. A frozen cutoff metric predicts the effective placement bracket in most held-out models and outperforms depth and value-copy baselines, while a one-layer shift in Qwen3-8B sharply reduces reach, showing that placement is sharp and not determined by remaining depth alone.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp