Nowait reasoning optimizer skill

Implements the NOWAIT technique for efficient reasoning in R1-style LLMs.

by davila7·MIT license·★ 32,299 Stars on the repo·GitHub ↗

Use now

Files of Nowait reasoning optimizer

davila7/main1 file shown
SKILL.md
Show the full text146 lines

NOWAIT Reasoning Optimizer

Implements the NOWAIT technique from the paper "Wait, We Don't Need to 'Wait'! Removing Thinking Tokens Improves Reasoning Efficiency" (Wang et al., 2025).

Overview

NOWAIT is a training-free inference-time intervention that suppresses self-reflection tokens (e.g., "Wait", "Hmm", "Alternatively") during generation, reducing chain-of-thought (CoT) trajectory length by 27-51% without compromising model utility.

When to Use

  • Deploying R1-style reasoning models with limited compute
  • Reducing inference latency for production systems
  • Optimizing token costs for reasoning tasks
  • Working with verbose CoT outputs that need streamlining

Supported Models

Model Series Type Token Reduction
QwQ-32B RL-based 16-31%
Phi4-Reasoning-Plus RL-based 23-28%
Qwen3-32B RL-based 13-16%
Kimi-VL-A3B Multimodal 40-60%
QvQ-72B-Preview Multimodal 20-30%

Important: NOWAIT works best with RL-based models. Distilled models (Qwen3-4B/8B/14B) show degraded performance when reflection tokens are suppressed.

Quick Start

1. Basic Implementation
from scripts.nowait_processor import NOWAITLogitProcessor

# Initialize processor for your model's tokenizer
processor = NOWAITLogitProcessor(tokenizer)

# Use during generation
outputs = model.generate(
    inputs,
    logits_processor=[processor],
    max_new_tokens=32768
)
2. Keywords Suppressed

See references/keywords.md for the complete list. Core keywords:

wait, alternatively, hmm, but, however, check, 
double-check, maybe, verify, again, oh, ah

How It Works

  1. Initialize Keywords: Identify reflection keywords from empirical analysis
  2. Expand to Token Variants: Map keywords to all token variants in vocabulary (e.g., "wait" → " wait", "Wait", " Wait", ".wait", "WAIT")
  3. Suppress During Inference: Set logits of reflection tokens to large negative values during decoding
Logits (Before)         Logits (After)
Wait     0.8     →     Wait     -inf
First    0.6     →     First    0.6
Hmm      0.5     →     Hmm      -inf
Let      0.4     →     Let      0.4

Key Findings

Why It Works
  • NOWAIT doesn't eliminate self-reflection entirely—it guides models to skip unnecessary "waiting" reasoning
  • Models still perform essential verification at key decision points
  • Results in more linear, straightforward reasoning paths
RL vs Distilled Models
Model Type NOWAIT Effect Recommendation
RL-based (QwQ, Phi4, Qwen3-32B) Stable accuracy, significant token reduction ✅ Recommended
Distilled (Qwen3-4B/8B/14B) Accuracy degradation on hard tasks ⚠️ Use with caution

Distilled models rely heavily on CoT structure from training data—removing reflection tokens disrupts their reasoning patterns.

Integration Examples

HuggingFace Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
from scripts.nowait_processor import NOWAITLogitProcessor

model = AutoModelForCausalLM.from_pretrained("Qwen/QwQ-32B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/QwQ-32B")

processor = NOWAITLogitProcessor(tokenizer)

response = model.generate(
    tokenizer(prompt, return_tensors="pt").input_ids,
    logits_processor=[processor],
    max_new_tokens=32768,
    do_sample=True,
    temperature=0.7
)
vLLM
from vllm import LLM, SamplingParams
from scripts.nowait_processor import get_nowait_bad_words_ids

llm = LLM(model="Qwen/QwQ-32B")
bad_words_ids = get_nowait_bad_words_ids(llm.get_tokenizer())

sampling_params = SamplingParams(
    max_tokens=32768,
    bad_words_ids=bad_words_ids
)

Expected Results

Task Type Original Tokens NOWAIT Tokens Reduction
Math (AIME) 15,000 10,500 30%
Visual QA (MMMU) 2,900 1,450 50%
Video QA (MMVU) 1,700 1,250 27%

Limitations

  • Less effective on very simple problems where CoT overhead is already minimal
  • Distilled models may suffer accuracy loss on challenging tasks
  • Some domains may require model-specific keyword tuning

References

  • Paper: arXiv:2506.08343v2
  • Complete keyword list: references/keywords.md
  • Implementation: scripts/nowait_processor.py
1---
2name: nowait-reasoning-optimizer
3description: Implements the NOWAIT technique for efficient reasoning in R1-style LLMs. Use when optimizing inference of reasoning models (QwQ, DeepSeek-R1, Phi4-Reasoning, Qwen3, Kimi-VL, QvQ), reducing chain-of-thought token usage by 27-51% while preserving accuracy. Triggers on "optimize reasoning", "reduce thinking tokens", "efficient inference", "suppress reflection tokens", or when working with verbose CoT outputs.
4---
5 
6# NOWAIT Reasoning Optimizer
7 
8Implements the NOWAIT technique from the paper "Wait, We Don't Need to 'Wait'! Removing Thinking Tokens Improves Reasoning Efficiency" (Wang et al., 2025).
9 
10## Overview
11 
12NOWAIT is a training-free inference-time intervention that suppresses self-reflection tokens (e.g., "Wait", "Hmm", "Alternatively") during generation, reducing chain-of-thought (CoT) trajectory length by **27-51%** without compromising model utility.
13 
14## When to Use
15 
16- Deploying R1-style reasoning models with limited compute
17- Reducing inference latency for production systems
18- Optimizing token costs for reasoning tasks
19- Working with verbose CoT outputs that need streamlining
20 
21## Supported Models
22 
23| Model Series | Type | Token Reduction |
24|--------------|------|-----------------|
25| QwQ-32B | RL-based | 16-31% |
26| Phi4-Reasoning-Plus | RL-based | 23-28% |
27| Qwen3-32B | RL-based | 13-16% |
28| Kimi-VL-A3B | Multimodal | 40-60% |
29| QvQ-72B-Preview | Multimodal | 20-30% |
30 
31**Important**: NOWAIT works best with RL-based models. Distilled models (Qwen3-4B/8B/14B) show degraded performance when reflection tokens are suppressed.
32 
33## Quick Start
34 
35### 1. Basic Implementation
36 
37```python
38from scripts.nowait_processor import NOWAITLogitProcessor
39 
40# Initialize processor for your model's tokenizer
41processor = NOWAITLogitProcessor(tokenizer)
42 
43# Use during generation
44outputs = model.generate(
45 inputs,
46 logits_processor=[processor],
47 max_new_tokens=32768
48)
49```
50 
51### 2. Keywords Suppressed
52 
53See `references/keywords.md` for the complete list. Core keywords:
54 
55```
56wait, alternatively, hmm, but, however, check,
57double-check, maybe, verify, again, oh, ah
58```
59 
60## How It Works
61 
621. **Initialize Keywords**: Identify reflection keywords from empirical analysis
632. **Expand to Token Variants**: Map keywords to all token variants in vocabulary (e.g., "wait" → " wait", "Wait", " Wait", ".wait", "WAIT")
643. **Suppress During Inference**: Set logits of reflection tokens to large negative values during decoding
65 
66```
67Logits (Before) Logits (After)
68Wait 0.8 → Wait -inf
69First 0.6 → First 0.6
70Hmm 0.5 → Hmm -inf
71Let 0.4 → Let 0.4
72```
73 
74## Key Findings
75 
76### Why It Works
77 
78- NOWAIT doesn't eliminate self-reflection entirely—it guides models to skip **unnecessary** "waiting" reasoning
79- Models still perform essential verification at key decision points
80- Results in more linear, straightforward reasoning paths
81 
82### RL vs Distilled Models
83 
84| Model Type | NOWAIT Effect | Recommendation |
85|------------|---------------|----------------|
86| RL-based (QwQ, Phi4, Qwen3-32B) | Stable accuracy, significant token reduction | ✅ Recommended |
87| Distilled (Qwen3-4B/8B/14B) | Accuracy degradation on hard tasks | ⚠️ Use with caution |
88 
89Distilled models rely heavily on CoT structure from training data—removing reflection tokens disrupts their reasoning patterns.
90 
91## Integration Examples
92 
93### HuggingFace Transformers
94 
95```python
96from transformers import AutoModelForCausalLM, AutoTokenizer
97from scripts.nowait_processor import NOWAITLogitProcessor
98 
99model = AutoModelForCausalLM.from_pretrained("Qwen/QwQ-32B")
100tokenizer = AutoTokenizer.from_pretrained("Qwen/QwQ-32B")
101 
102processor = NOWAITLogitProcessor(tokenizer)
103 
104response = model.generate(
105 tokenizer(prompt, return_tensors="pt").input_ids,
106 logits_processor=[processor],
107 max_new_tokens=32768,
108 do_sample=True,
109 temperature=0.7
110)
111```
112 
113### vLLM
114 
115```python
116from vllm import LLM, SamplingParams
117from scripts.nowait_processor import get_nowait_bad_words_ids
118 
119llm = LLM(model="Qwen/QwQ-32B")
120bad_words_ids = get_nowait_bad_words_ids(llm.get_tokenizer())
121 
122sampling_params = SamplingParams(
123 max_tokens=32768,
124 bad_words_ids=bad_words_ids
125)
126```
127 
128## Expected Results
129 
130| Task Type | Original Tokens | NOWAIT Tokens | Reduction |
131|-----------|-----------------|---------------|-----------|
132| Math (AIME) | 15,000 | 10,500 | 30% |
133| Visual QA (MMMU) | 2,900 | 1,450 | 50% |
134| Video QA (MMVU) | 1,700 | 1,250 | 27% |
135 
136## Limitations
137 
138- Less effective on very simple problems where CoT overhead is already minimal
139- Distilled models may suffer accuracy loss on challenging tasks
140- Some domains may require model-specific keyword tuning
141 
142## References
143 
144- Paper: arXiv:2506.08343v2
145- Complete keyword list: `references/keywords.md`
146- Implementation: `scripts/nowait_processor.py`

Discussion

Alternatives

AI engineerAct as an expert AI engineer specializing in practical machine learning implementation and AI integration for production applications, ensuring efficient and robust AI solutions.Data & AI · CC0-1.0OneKGPd: Individual-Level Queries over the 1000 Genomes ProjectQuery the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants. Use when a question is about individuals or variants in the 1000 Genomes Project cohort: which individuals carry variants matching specific criteria in a gene or region, which individuals are homozygous-reference at a position, which variants exist in the dataset or carried by specified individuals in a gene or region, the relatedness between two specified individuals. Variants are returned with 1000 Genomes allele frequencies (AF), gnomAD v4.1 exome and genome AF, AlphaMissense score, and HGVSp annotations.Science · MITPyMC Bayesian ModelingBayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.Science · MITStatsmodels: Statistical Modeling and EconometricsStatistical models library for Python. Use when you need specific model classes (OLS, GLM, mixed models, ARIMA) with detailed diagnostics, residuals, and inference. Best for econometrics, time series, rigorous inference with coefficient tables. For guided statistical test selection with APA reporting use statistical-analysis.Science · MIT