Program-as-Weights: A Programming Paradigm for Fuzzy Functions
AuthorsWentao Zhang, Liliana Hotsko, Woojeong Kim, Pengyu Nie, Stuart Shieber, Yuntian Deng
Resources
Program-as-Weights lets a model compile a natural-language task description into a small reusable neural program that can run locally and cheaply instead of calling a giant LLM every time.
Key results
dataset used to train the PAW compiler
text-to-LoRA default exact match at rank 64
precursor exact match on FuzzyBench
Qwen3 0.6B interpreter executing PAW programs on verified FuzzyBench test set
direct prompting baseline on FuzzyBench
tokens per second on a MacBook M3
What the paper found
Program-as-Weights (PAW), from researchers at the University of Waterloo, Cornell University, and Harvard University, reframes fuzzy software tasks as compile-once, run-locally neural programs. A natural-language specification is first converted into a discrete pseudo-program by an off-the-shelf Qwen3-4B-Instruct-2507 compiler, then into a per-function PEFT artifact by a trained 4B Qwen3 LoRA compiler; the resulting program is executed by a frozen Qwen3 0.6B interpreter. Trained on FuzzyBench, a 10M-example dataset spanning 29 thematic versions and 800+ task categories, PAW’s default LoRA instantiation reaches 65.7% exact match with rank 64, versus 50.4% for prefix-tuning and 9.8% for direct prompting. On the verified FuzzyBench test set, the Qwen3 0.6B PAW system achieves 73.78% exact match, surpassing direct prompting of Qwen3-32B at 68.70% while using about 50× less inference memory, and it runs at 30 tokens/s on a MacBook M3 with a ∼430 MB shared base plus a 23 MB per-program adapter. The paper also shows a multimodal extension using Qwen3-VL-4B as compiler, where PAW LoRA reaches 0.274, 0.414, and 0.552 on Circuit, Chemical, and Music diagram tasks, respectively, and reports only minor degradation under heavy noisy-specification perturbations because the pseudo-program denoises the input before execution.
Original abstract
Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking search results by intent, and are increasingly outsourced to large language model APIs at the cost of locality, reproducibility, and price. We propose fuzzy-function programming: compiling such a function from a natural-language specification into a compact, locally-executable neural artifact. We instantiate this paradigm with Program-as-Weights (PAW), in which a 4B compiler trained on FuzzyBench, a 10M-example dataset we release, emits parameter-efficient adapters for a frozen, lightweight interpreter. A 0.6B Qwen3 interpreter executing PAW programs matches the performance of direct prompting of Qwen3-32B, while using roughly one fiftieth of the inference memory and running at 30 tokens/s on a MacBook M3. PAW reframes the foundation model from a per-input problem solver into a tool builder: invoked once per function definition, it produces a small reusable artifact whose subsequent calls per function application are cheap and offline.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.