Agents

Laya Setup Guide: A Free Local Decision Engine for Claude Code

7 minute readUpdated October 2026Explore more

TL;DR

Laya is a free, open-source model that makes quick, typed decisions instead of writing text. You give it some text and a question like pick a label, score it, or yes or no, and it answers in one pass with a confidence number. It runs on your own computer, works in 100+ languages, and plugs into Claude Code as an MCP server. It is fast on a GPU and slower on a normal CPU, and it gets much more accurate once you fine-tune it on your own decisions.

Learn Claude Code. Earn income. Only $9.

👉 https://www.skool.com/claudecodeclub

What Laya is

NandhaKishorM/laya

Over 30,000 stars on October 2, 2026. Apache 2.0 license. Models on Hugging Face: https://huggingface.co/convaiinnovations/laya

Most AI models write text. Laya does not. It is a "System 1" decision engine: you hand it some text (an email, a ticket, a document) and a set of typed questions, and it returns the answers with a calibrated confidence for each, in a single forward pass. Because it never generates text, there is nothing to parse and nothing to make up.

It has three question types:

  • choice: pick one label, like which department should handle this ticket.
  • score: rate on a scale, like how urgent is this: not urgent, soon or blocking.
  • noul: yes or no, returned as the probability that the answer is yes, like does this customer threaten to cancel.

It works in 100+ languages, and a built-in router picks the right model for each request. The README measures 33 ms for one question and 7.2 ms per question when batched, on a T4 GPU.

What you need

  • Claude Code installed.
  • Python 3.10 or newer.
  • A normal computer works: Laya runs on CPU, Apple Silicon (MPS) or an NVIDIA GPU. A GPU is just faster.
  • Internet the first time you run a prediction, so it can download the model from Hugging Face.

Step 1: Install Laya

Easiest: paste this into Claude Code:

promptRead https://github.com/NandhaKishorM/laya and install Laya for me in a Python virtual environment in this folder. Then run one test prediction and show me the result.

Or do it yourself:

bashpython -m pip install laya

Then try the built-in command line. The first line only routes and works offline. The second gives full answers from a ready-made triage question set (it downloads the model the first time):

bashlaya "I was charged twice, please refund"
laya "My payment failed twice" --preset triage

Step 2: Connect Laya to Claude Code (MCP)

Laya ships an optional MCP server, so Claude Code can call it as a tool for typed decisions. Install the extra:

bashpip install "laya[mcp]"

Then add it to Claude Code. The LAYA_DEVICE=cpu setting is the README's own example; drop it to let Laya pick your GPU automatically:

bashclaude mcp add laya --env LAYA_DEVICE=cpu -- laya-mcp-server

Restart Claude Code and ask it to use Laya:

promptUse the Laya tools to sort these 10 support emails. For each one, decide the department (billing, technical, other), the urgency (not urgent, soon, blocking) and whether the customer threatens to cancel. Show me a table with the confidence for each answer.

Good things to use it for

  • Support triage: route tickets by department, urgency and churn risk.
  • Inbox sorting: lead, spam or needs a human, in any of 100+ languages.
  • Prompt guardrails: flag jailbreak or injection attempts before they reach your agent.
  • Model routing: decide whether a request needs a big model or a small one, so you spend less on easy tasks.

Make it accurate: fine-tune on your own decisions

Out of the box Laya works zero-shot, but the README is clear that fine-tuning on your own domain is where accuracy jumps. On their typed-decisions benchmark of 2,000 decisions, the fine-tuned model scored 0.766 accuracy against 0.362 for the base English model. The repo includes a Kaggle notebook and an Apple Silicon script that run the whole loop.

The catches

  • 33 ms is on a GPU. That number was measured on a T4 GPU. On a normal CPU it is slower: the README's own table shows a few hundred milliseconds per request.
  • Zero-shot accuracy can be low on your domain. Until you fine-tune on your own decisions, check its answers carefully.
  • Long inputs get less reliable. In the README's test, results stayed strong up to about 4,000 tokens of text and varied beyond that.
  • Test on your own data before automating anything that matters. The confidence scores help, but pick your thresholds from results on your real data.

Quick start

  1. 1pip install laya and run the triage preset to see it work.
  2. 2Add the MCP server to Claude Code with claude mcp add.
  3. 3Ask Claude Code to use Laya on a small batch of your real emails or tickets.
  4. 4Check the answers by hand before you trust it with anything automatic.
Learn Claude Code. Earn income. Only $9.

👉 https://www.skool.com/claudecodeclub

Common questions

  • Is Laya free?

    Yes. Laya is open source under the Apache 2.0 license and runs on your own computer, so there is no per-request fee.

  • Do I need a GPU to run Laya?

    No. It runs on CPU and Apple Silicon too. The 33 ms speed was measured on a T4 GPU, so expect it to be slower on a normal CPU.

  • Can Laya replace Claude?

    No. Laya only answers typed questions: pick a label, give a score or yes or no. It does not write text or answer open questions. It works best as a fast helper that Claude Code calls.

  • How accurate is Laya?

    It depends on your data. Zero-shot results can be low on your domain. On the repo's benchmark, fine-tuning lifted accuracy from 0.362 to 0.766, so test it on your own decisions first.

Get more free tools like this inside Claude Code Club for $9/mo

Get 650+ plug-and-play skills, MCPs & prompts, plus 8,000+ members - $9/mo, cancel anytime.

Join the Club