When Gnosys saves tokens.
And when it does not.
Our estimate: Gnosys can save tokens when it replaces a large, always-loaded memory file and you retrieve only what a task needs. Below break-even, it likely adds context. That is about 5k tokens with core loaded eagerly, or about 4k with verified tool deferral. Under about 2.5k, it adds context even with deferral and a single retrieval.
Selective memory still has a fixed cost.
The core toolset is about 3.7k tokens before retrieval, according to our docs. In the model below, one or two targeted retrievals put break-even at about 5.2–6.7k tokens of replaced memory. A small index with selective file reads is already a lean baseline. [1]
M tokens
S + k × r tokens
Compare the memory you actually load today. A full conversation transcript is a different baseline. This model compares context size, not total billed usage.
The tool schemas count too.
Tool counts are documented and verified in source. Payload sizes come from our docs. Token figures are approximations at about 4 characters per token, not tokenizer measurements. This is the fixed cost when a client loads all schemas in the selected tier. [1]
| Starting tier | Tools | Payload | Approx. tokens |
|---|---|---|---|
| core (default) | 19 | ~15k chars | ~3.7k |
| standard | 34 | ~25k chars | ~6.3k |
| full | 56 | ~41k chars | ~10.2k |
Deferred tool loading can lower the starting cost. Verify what your client actually loads. The deferred scenario below is an estimate, not a guarantee for Claude Desktop.
Does your memory exceed break-even?
Gnosys uses less memory context in this simplified model when M is greater than S + k × r. We assume r ≈ 1.5k tokens for a retrieval plus a follow-up read. Per-call retrieval sizes use a rough 4 characters per token heuristic. They are not measured usage.
M > S + k × r
- M
- Always-loaded memory replaced
- S
- Fixed tool-schema cost
- k
- Retrievals per session
- r
- Tokens per retrieval + read
| Scenario | S | k | S + k × r | Saves context if M exceeds about |
|---|---|---|---|---|
| Core, light use | 3.7k | 1 | ~5.2k | ~5k tokens |
| Core, typical use | 3.7k | 2 | ~6.7k | ~7k tokens |
| Core, heavy use | 3.7k | 4 | ~9.7k | ~10k tokens |
| Deferred tools, typical use | ~1.0k | 2 | ~4.0k | ~4k tokens |
| Full, typical use | 10.2k | 2 | ~13.2k | ~13k tokens |
This model leaves out tool-call output, repeated input across model turns, cache reads and writes, memory writes, and any external provider calls. Include those costs in your A/B run. No dollar savings are claimed here.
Your baseline decides the result.
Likely to save tokens
- You replace a memory file larger than about 5–7k tokens and make one or two targeted retrievals.
- Most tasks use a small part of your stored memory.
- You keep the starting toolset small and limit repeated searches.
Likely to add tokens
- Your always-loaded memory is below break-even: about 5k tokens with core loaded eagerly, or about 4k with verified deferral. Under about 2.5k, it adds context in every scenario here.
- You already use a small index with selective file reads.
- The agent searches repeatedly, retrieves duplicates, or reads long memories in full.
- You load the full tier for a task that only needs core retrieval.
Recall defaults are verified in source: maxMemories: 8, aggressive: true, and minRelevance: 0.4. Recall formatting truncates each body snippet to 250 characters. Aggressive mode keeps the top three results before applying the relevance floor to the rest. These limits bound snippets, not every possible retrieval. gnosys_read returns a full memory. [2]
Estimated search-type output is about 0.6–1.4k tokens per call; an atomic gnosys_read may add about 150–600 tokens. Titles, metadata, result count, and body length change those sizes. A long read can exceed that range.
Start small. Retrieve with a purpose.
Start on core and keep the tool list stable.
Set the starting tier in your MCP server environment. Agents can still switch with gnosys_toolset, so instruct them to keep core for the session. If a task needs a larger tier, choose it at startup. [1]
"env": { "GNOSYS_MCP_TOOLSET": "core" }
Avoid mid-session escalation. Tool-list changes can invalidate a client's prompt cache. Verify cache behavior in your own usage logs.
Verify deferral and keep the pointer file lean.
Use deferred tools where your client supports them. Check the loaded schemas rather than assuming deferral works. Keep CLAUDE.md to a short pointer with identity, scope, and retrieval instructions. Review generated sync blocks so you do not load the same memory twice.
Discover first. Read only what you need.
Use gnosys_discover for metadata, with a small result limit such as 5–8. Then use gnosys_read for the one or two memories the task needs. Keep memories atomic and remove stale duplicates through your normal maintenance workflow.
Compare equal tasks and equal quality.
Use 10–20 representative tasks with identical prompts, model, and settings. Record the client version, starting tier, and whether tool deferral is active. Control session length and cache conditions in both arms.
| Arm | Memory setup | Purpose |
|---|---|---|
| A · Baseline | Current markdown memory, no Gnosys | Measure what you use today |
| B · Core | Lean pointer file + Gnosys core | Include schema and retrieval overhead |
| C · Optional | Same as B, with verified tool deferral | Measure your client's deferred path |
Record starting context.
In Claude Code, inspect /context for tool schemas and memory-file usage. For exact schema counts, use your model provider's token-counting tool.
Record the complete session.
Capture input, cache creation, cache read, and output tokens. Count memory calls and repeated searches. Include any separate provider usage for synthesis, embeddings, or maintenance. Record elapsed time too.
Check the answer, then compare.
Score correct recall and completed tasks. Compare median total usage at equal or better quality. For a cost comparison, apply your provider's current rates to each token category. A smaller starting context alone does not prove a saving.
What is verified. What is estimated.
Verified here means checked against our product docs or source, not independently benchmarked. Retrieval sizes, deferred costs, and break-even thresholds are estimates. This page does not provide an independent Gnosys token-usage benchmark.
Use your client and provider usage records. Gnosys does not enforce a spend budget. Configured synthesis and LLM-backed maintenance can add provider costs. [3]
The phrase "Zero Context Bloat" is not a universal guarantee. Tool schemas and retrieved content use context. Client deferral, caching, repeated calls, and task quality determine the measured result.
- MCP toolset tiers. Our documentation for tier counts, approximate schema sizes, starting-tier configuration, and runtime switching.
- Recall source. Verified defaults and the 250-character snippet limit. Product documentation describes retrieval tools.
- Cost and limits. No cumulative token or dollar tracking, and no enforced spend budget.