HN
Today

Build your own decision model

This technical deep dive explores how to build 'System one' decision models using Large Language Models by constraining their output for quick, single-pass responses. It provides practical code examples demonstrating constrained decoding and tackles the crucial issue of model calibration. The article resonates on HN for its hands-on approach to enhancing LLM reliability in specific decision-making applications.

11
Score
1
Comments
#1
Highest Rank
22h
on Front Page
First Seen
Oct 10, 11:00 PM
Last Seen
Oct 11, 8:00 PM
Rank Over Time
212222221223610810141316202425

The Lowdown

The article introduces 'System one' decision models, which are designed to infer and respond with calibrated probabilities for a fixed set of answers. Unlike general LLM output generation that tokenizes sequentially, these models aim for single-pass decision-making by constraining the output to specific options.

  • The author illustrates this concept by contrasting standard LLM generation, which can require multiple passes for structured output, with constrained decoding for decision models, which processes input in a single pass by masking irrelevant vocabulary tokens.
  • A practical Python implementation is provided, using a Qwen/Qwen3-1.7B model from the transformers library to demonstrate how to constrain LLM output to a predefined set of answer options.
  • The example shows the model answering a simple question, 'What color is the sky?', with high confidence in the correct answer ('Blue'), and includes initial accuracy metrics from running the model against a CommonsenseQA dataset.
  • A significant challenge highlighted is model miscalibration, where the LLM can be overconfident in its predictions. This is demonstrated with an ambiguous question about where to find a 'bat,' where the model assigns near-certainty to 'Cave' despite other plausible interpretations.
  • To address miscalibration, the article explains and demonstrates temperature scaling, a technique used to adjust the model's output probability distribution so that its confidence scores more accurately reflect its true accuracy.
  • The author shares a GitHub repository with scripts for building, evaluating, finetuning, and calibrating one's own decision model, encouraging readers to experiment with different models.

By leveraging constrained decoding and meticulous calibration, the article offers a blueprint for transforming general-purpose LLMs into more reliable and interpretable 'System one' decision-making tools, crucial for applications requiring high confidence and accuracy.