Token-Oriented Object Notation is a row-based format that declares keys once and repeats only values. It is not a new idea — it is the oldest trick in database storage applied to text serialisation.
Every time you send a JSON array to an LLM, you pay the same tax twice: once for the actual data, and once for the wrapper. Every key in {"name":"Alice","age":32,"role":"admin"} gets re-emitted for every single row. In a thousand-row dataset, the word "name" appears a thousand times. The LLM pays for every token.
TOON separates the schema from the data. Keys are declared once at the top. Rows follow. The parser — or the model — reconstructs objects on the fly. The result is a format with significantly fewer tokens for arrays of homogeneous objects, without losing structural meaning.
Here is a small user list in both formats side by side.
[
{ "id": 1, "name": "Alice", "role": "admin", "active": true },
{ "id": 2, "name": "Bob", "role": "viewer", "active": true },
{ "id": 3, "name": "Carol", "role": "editor", "active": false }
]
# keys declared once
@keys id name role active
# one value row per object
1 Alice admin true
2 Bob viewer true
3 Carol editor false
→ Try this example in ilovejson TOON converter ↗
The keys disappear from every row. A TOON parser zips the declared header against each row the same way Python's zip() or SQL's SELECT column list works. For three rows the saving looks modest. For 500 rows the JSON version is around 50 KB; the TOON version sits under 8 KB.
Paste any TOON-formatted data into the ilovejson TOON to JSON converter to instantly see the reconstructed JSON objects, explore the tree view, and validate the output. No sign-up, no data upload — runs entirely in your browser.
The first objection to any columnar format is always nesting. Real-world data is not flat. A user might have an embedded address or a list of tags. TOON handles this with dot-notation keys and an inline JSON escape for anything genuinely complex.
@keys id name address.city address.country tags
1 Alice London UK ["admin","billing"]
2 Bob Berlin DE ["viewer"]
3 Carol Toronto CA []
→ Try nested TOON example in ilovejson ↗
Dot-notation keys are expanded into nested objects during parsing. So address.city becomes { "address": { "city": "..." } }. For values that cannot be flattened — deep nested arrays, polymorphic objects — the value is a JSON fragment sitting inline. The parser detects the opening [ or { and hands that token off to a JSON sub-parser before continuing along the row.
This is not a hack. CSV does exactly the same thing when a cell contains a quoted comma. Every columnar format needs an escape hatch for values that contain the delimiter. TOON's escape hatch is JSON itself, which means the parser is already available on every platform that speaks JSON.
Consider an analytics event stream — the kind of data that gets sent to an LLM for summarisation or fed into a fine-tuning pipeline.
[
{
"event": "page_view",
"ts": 1720000000,
"user": { "id": "u1", "plan": "pro" },
"meta": { "path": "/dashboard", "ref": "google" }
},
{
"event": "click",
"ts": 1720000045,
"user": { "id": "u1", "plan": "pro" },
"meta": { "path": "/dashboard", "ref": "null" }
}
]
@keys event ts user.id user.plan meta.path meta.ref
page_view 1720000000 u1 pro /dashboard google
click 1720000045 u1 pro /dashboard null
→ Try this example in ilovejson TOON converter ↗
The second version is roughly a third of the character count. The structure is identical once reconstructed. Keys that are constant across all rows — user.id, user.plan here — appear only once in the header. Any repeated context that JSON would duplicate on every line is simply gone.
The common worry about columnar formats is debuggability. JSON is self-describing: you can drop any fragment into a viewer and understand it immediately. TOON requires the header to decode any row. Lose the header, lose the schema.
In practice this is manageable, and the tooling concern is more real than the format concern. A few things worth knowing:
Column count mismatches are immediately visible. If a row has five values and the header has six keys, every TOON parser raises an error on that exact line. JSON silently omits the field or returns undefined depending on the consumer. TOON's strictness is a feature.
Headers can be embedded at regular intervals. For streams and log files, you can repeat the @keys declaration every N rows — similar to CSV files that repeat the header every thousand lines for easier slicing with tools like grep or head. The parser skips repeated headers.
Any JSON viewer that accepts a transformation step can render TOON. The ilovejson formatter accepts a pre-processor: paste TOON, click Parse as TOON (once the parser is wired), and the tree view expands the reconstructed objects exactly as it would for JSON. The raw text view is always available if you need to inspect the format directly.
There is no official RFC for TOON. It is a pattern, not a standard — and that is fine. Every language has the primitives needed to implement a basic parser in under 50 lines.
def parse_toon(text):
lines = [l for l in text.strip().splitlines() if l and not l.startswith("#")]
keys = lines[0].split()[1:] # strip @keys prefix
rows = []
for line in lines[1:]:
values = line.split(None, len(keys) - 1)
rows.append(dict(zip(keys, values)))
return rows
For production use, swap str.split() with a tokeniser that respects quoted fields and inline JSON fragments. The csv module in Python's standard library can handle quoted fields; treat it as a CSV parser with a space delimiter and you are 90% there. A pickle serialisation step on top gives you a binary TOON store for caching — same schema, binary values, no parsing overhead on repeated loads.
def parse_toon(text)
lines = text.lines.map(&:strip).reject { |l| l.empty? || l.start_with?("#") }
keys = lines.first.split.drop(1) # drop @keys
lines.drop(1).map { |row| keys.zip(row.split(nil, keys.size)).to_h }
end
A gem does not exist yet. That is actually the right time to write one. The format is simple enough that a gem for TOON — with a proper tokeniser, a streaming reader, and a JSON-roundtrip helper — is a realistic weekend project. The interface would mirror Ruby's CSV standard library: TOON.parse(str), TOON.generate(rows), TOON.foreach(file) { |row| }.
function parseToon(text) {
const lines = text.split("\n")
.map(l => l.trim())
.filter(l => l && !l.startsWith("#"));
const keys = lines[0].split(" ").slice(1);
return lines.slice(1).map(row => {
const vals = row.split(" ", keys.length);
return Object.fromEntries(keys.map((k, i) => [k, vals[i]]));
});
}
An npm package follows the same shape. Publish it as toon-parse, export a parse() and stringify(), add a streaming interface for large files, and you have a format library that competes with csv-parse in the structured-data-for-LLMs niche.
Accept: text/toon. Return TOON when the header is present, JSON otherwise. Measure the token count difference on the LLM side. The format pays for itself on the first call that is long enough to hit a context limit.
Token overhead is not the only metric that matters, but it is the metric that directly maps to API cost when feeding data to language models.
| Format | Token Overhead | Best Used For |
|---|---|---|
| CSV / TSV | Lowest — no structural tokens | Purely flat tables or sheets |
| TOON | Extremely low — declares keys once | Arrays of structured objects, AI prompts |
| YAML | Low — removes braces and quotes | Hierarchical configs and single objects |
| JSON | High — repeats keys, syntax-heavy | API responses and strict type generation |
TOON is not a replacement for JSON. It solves one specific problem — token waste in homogeneous arrays — and solves it well. It is the wrong choice for:
CSV is still the lowest-overhead format for flat data. If your rows have no nesting and no optional fields, CSV beats TOON on token count and beats it on tooling. Every language has a CSV parser. Not every language has a TOON parser — yet.
TOON earns its place exactly in the middle: data that is too structured for CSV (nested keys, optional inline JSON fragments) and too repetitive for JSON (hundreds of rows where every key repeats). That is the shape of most LLM batch inputs in practice — analytics events, user records, product catalogues, log lines. TOON is a pragmatic answer to that shape.
Whether it becomes a gem, a pip package, or an npm module depends on whether developers find the pattern useful enough to reach for it. The implementation is short. The payoff on long prompts is real. The rest is just momentum.
Try pasting a TOON-formatted dataset into ilovejson — format it as JSON to explore the reconstructed object tree.