Always-on agents that live on your own computer.

Tholos is an open-source app that gives a small team of AI agents one shared workspace: tables, notes and a task board. They work on a schedule, pass jobs to each other and pause for your approval where your rules require it. The default model is 1.6 GB and runs on an ordinary CPU.

On a CPU, Tholos-2B passes 134 of the 160 Tholos-Bench scenarios, 17 more than its base model. Built by senior AI engineer Mert Kaya.

Tholos at work

The Tholos board: Scout and Writer on their schedules, Scout at work, nothing waiting for a decision.
The board during a real run on a CPU: Scout collects new papers and posts into a shared leads table.

How it works

Give each agent a job.

An agent is a name, a role, a model, a schedule and a short list of tools. Its jobs are plain sentences, such as “Every two hours, read the Sources note, check each site and add anything new to the leads table.”

One workspace for the whole team.

Agents read and write the same tables and notes. Each change records who made it and which version it replaced. When two agents touch the same row, the second write is refused until that agent has read the new value.

You keep the keys.

Rules decide what an agent may do alone, what waits for your approval and what it may never do. By default, fetching a web page asks you first. When an agent needs you, it pauses with the exact action on screen: the web address it wants to open, the rows it wants to change. Approve it once, deny it or always allow it, in the browser or, once connected, on Telegram.

The model

Tholos-2B is MiniCPM5-2B fine-tuned on synthetic agent runs, each kept only if its final workspace passed the scenario's checks. The scores below come from a Kaggle T4 GPU. The technical report covers training and evaluation.

Tholos-Bench: 160 scenarios, each scored on the final state of the workspace.
ModelFile sizePassed (of 160)Injection passed (of 12)
Qwen3.5-2B1.28 GB974
MiniCPM5-2B (base)1.56 GB1128
Tholos-2B this model1.56 GB13710
Llama 3.2 3B Instruct2.02 GB371
Granite 4.2-3B2.24 GB1297
Qwen3.5-4B2.74 GB1416

Every model ran as a Q4_K_M file on one Kaggle T4 GPU with llama.cpp, JSON schema decoding, temperature 0 and the same 16,384-token context. The injection column covers the 12 scenarios where a page, a table cell or a note carries instructions the agent should ignore. Only Tholos-2B was fine-tuned for the Tholos step format, one JSON step per turn. LFM2.5-2.6B returned empty replies under JSON schema decoding and has no score.

Qwen3.5-4B passes 141 with a file 1.75 times the size. On the 12 injection scenarios Tholos-2B passes 10, the highest count in the table.

Tholos-2B is meant for English. Tholos also works with any model behind an OpenAI-compatible API, local or hosted, chosen per agent.

Install

You need Ollama and uv.

ollama pull hf.co/mertkayacs/Tholos-2B-GGUF:Q4_K_M
uv tool install git+https://github.com/mertkayacs/tholos
tholos

The last command starts Tholos at http://127.0.0.1:7070. Open it, then:

  1. In Settings, press Detect and add the model it finds.
  2. Load the Research desk team from Starter teams.
  3. Choose that model for its three agents: Scout, Analyst and Writer.

Scout checks Hacker News, arXiv and Hugging Face Papers every two hours, Analyst scores what it finds and Writer sends you a brief on Friday afternoon.