Give each agent a job.
An agent is a name, a role, a model, a schedule and a short list of tools. Its jobs are plain sentences, such as “Every two hours, read the Sources note, check each site and add anything new to the leads table.”
Tholos is an open-source app that gives a small team of AI agents one shared workspace: tables, notes and a task board. They work on a schedule, pass jobs to each other and pause for your approval where your rules require it. The default model is 1.6 GB and runs on an ordinary CPU.
On a CPU, Tholos-2B passes 134 of the 160 Tholos-Bench scenarios, 17 more than its base model. Built by senior AI engineer Mert Kaya.

An agent is a name, a role, a model, a schedule and a short list of tools. Its jobs are plain sentences, such as “Every two hours, read the Sources note, check each site and add anything new to the leads table.”
Agents read and write the same tables and notes. Each change records who made it and which version it replaced. When two agents touch the same row, the second write is refused until that agent has read the new value.
Rules decide what an agent may do alone, what waits for your approval and what it may never do. By default, fetching a web page asks you first. When an agent needs you, it pauses with the exact action on screen: the web address it wants to open, the rows it wants to change. Approve it once, deny it or always allow it, in the browser or, once connected, on Telegram.
Tholos-2B is MiniCPM5-2B fine-tuned on synthetic agent runs, each kept only if its final workspace passed the scenario's checks. The scores below come from a Kaggle T4 GPU. The technical report covers training and evaluation.
| Model | File size | Passed (of 160) | Injection passed (of 12) |
|---|---|---|---|
| Qwen3.5-2B | 1.28 GB | 97 | 4 |
| MiniCPM5-2B (base) | 1.56 GB | 112 | 8 |
| Tholos-2B this model | 1.56 GB | 137 | 10 |
| Llama 3.2 3B Instruct | 2.02 GB | 37 | 1 |
| Granite 4.2-3B | 2.24 GB | 129 | 7 |
| Qwen3.5-4B | 2.74 GB | 141 | 6 |
Every model ran as a Q4_K_M file on one Kaggle T4 GPU with llama.cpp, JSON schema decoding, temperature 0 and the same 16,384-token context. The injection column covers the 12 scenarios where a page, a table cell or a note carries instructions the agent should ignore. Only Tholos-2B was fine-tuned for the Tholos step format, one JSON step per turn. LFM2.5-2.6B returned empty replies under JSON schema decoding and has no score.
Qwen3.5-4B passes 141 with a file 1.75 times the size. On the 12 injection scenarios Tholos-2B passes 10, the highest count in the table.
Tholos-2B is meant for English. Tholos also works with any model behind an OpenAI-compatible API, local or hosted, chosen per agent.
ollama pull hf.co/mertkayacs/Tholos-2B-GGUF:Q4_K_M
uv tool install git+https://github.com/mertkayacs/tholos
tholosThe last command starts Tholos at http://127.0.0.1:7070. Open it, then:
Scout checks Hacker News, arXiv and Hugging Face Papers every two hours, Analyst scores what it finds and Writer sends you a brief on Friday afternoon.