REFERENCE / A LIVING TOOLGLASS DICTIONARY

The Glassary.

A Toolglass Dictionary of AI

Artificial intelligence has acquired an unfortunate habit of producing terminology almost as quickly as it produces text.

Some of it is useful. Some of it is marketing fog. Some of it names something genuinely new using language that makes you want to leave the building.

The Glassary is Toolglass's attempt to do better. Where a good term already exists, we keep it. Where language is muddy, we clarify it. Where something keeps happening but nobody has given it a useful name, we reserve the right to manufacture one.

THE TEST

Does the term help us notice, discuss or prevent something real?

If not, out it goes.
Entries
51
◈ Toolglass
42
↗ Common
6
⌁ Research
3
AINDEX ↑

AI Cowboy

◈

A person or organisation using powerful AI with high autonomy and low operational discipline.

The defining characteristic is not experimentation or risk-taking. It is allowing capability to expand faster than evidence, authority, reproducibility and verification.

An AI cowboy may produce excellent work. The problem is that nobody can reliably explain why the work should be trusted.

AI cowboying optimises for motion. AI engineering optimises for trustworthy progress.

SEE ALSOCowboy Compute · Frontier Code · Authority Fog · Proof Debt

Authority Bleed

◈

When permission granted for one operation, tool or scope quietly spreads beyond its intended boundary.

An agent authorised to edit one directory ends up capable of rewriting the repository. An approval to deploy one component becomes interpreted as approval to deploy everything.

Authority should travel deliberately, not seep through the floorboards.

SEE ALSOAuthority Fog · Autonomy Cliff · Blast Radius

Authority Fog

◈

A situation in which nobody can state precisely what an AI is permitted to change, approve, publish, merge, delete or deploy.

Everything may appear to be working perfectly until somebody asks, "Who actually authorised that?"

SEE ALSOAuthority Bleed · Human Hinge

Autonomy Cliff

◈

A point at which a seemingly small increase in AI freedom produces a disproportionately large increase in possible consequences.

Giving an agent one more tool can sometimes be equivalent to adding an entire new category of action.

SEE ALSOBlast Radius · Authority Bleed

BINDEX ↑

Blast Radius

↗

The maximum damage or disruption possible if a system, agent or decision goes wrong.

Borrowed from security and infrastructure engineering, the term becomes especially useful when AI systems can operate tools rather than merely generate text.

Good autonomous systems do not merely try to prevent failure. They also limit how much any single failure can affect.

Branch Ghost

◈

An obsolete or abandoned software branch that still looks authoritative enough to fool a human or agent into treating it as current work.

Branch ghosts become particularly troublesome in long-running AI projects where agents repeatedly rediscover old work without knowing its status.

SEE ALSOContext Archaeology · State Spill

Branch Necromancy

◈

The accidental resurrection of obsolete code, architecture or decisions because their relationship to current work was not understood.

A branch ghost is the corpse.

Branch necromancy is when somebody gets it walking again.

CINDEX ↑

CI Pilgrimage

◈

Sending work to hosted continuous integration out of habit when the important question could have been answered locally faster and more cheaply.

Hosted CI is valuable evidence. It should not become a sacred mountain that every tiny change must climb.

SEE ALSOCowboy Compute · Verification Theatre

Confidence Lacquer

◈

Beautifully structured and authoritative language applied over reasoning that is considerably weaker underneath.

The prose is polished. The headings are excellent. Unfortunately, nobody checked whether paragraph three was true.

SEE ALSOSloppy Slop · Evidence Laundering

Context Archaeology

◈

The work of reconstructing project history by excavating old conversations, branches, commits, logs, documents and forgotten directories.

Some archaeology is inevitable.

If every new agent needs a shovel, the project has a state-management problem.

SEE ALSOContext Debt · Context Sediment · Recovery Surface

Context Debt

◈

Future work created when important state, reasoning, evidence or decisions are not captured clearly when they occur.

Like technical debt, context debt feels cheap at the moment it is created.

Someone else pays later.

Frequently that someone is an AI burning thousands of tokens rediscovering Tuesday.

Context Engineering

↗

The deliberate management of the information available to an AI system during a task.

This includes far more than writing a prompt. It may involve tool results, memory, project state, retrieval, summaries, permissions, histories and what information is deliberately excluded.

Context engineering asks:

What does the system need to know now, and what should not be occupying its attention?

Context Rot

↗

The degradation of an AI system's performance as its context becomes stale, contradictory, bloated or increasingly irrelevant.

Information can remain technically present while becoming operationally harmful.

SEE ALSOContext Sediment · Context Debt

Context Sediment

◈

Historical information that gradually accumulates in an AI's working context.

Some sediment contains valuable history.

Some is simply mud.

Unlike context rot, sediment is not necessarily harmful. It becomes a problem when nobody distinguishes geological record from rubbish.

Cowboy Compute

◈

Compute wasted because nobody first established what was already known, tested, decided or available.

Agents reread repositories, repeat investigations, reproduce tests and reconstruct decisions that already existed somewhere else.

The machines appear extremely busy.

Progress remains strangely stationary.

SEE ALSOInference Bonfire · CI Pilgrimage · Context Debt

EINDEX ↑

Evidence Laundering

◈

The process by which weak, uncertain or unsupported information acquires apparent credibility through repetition or repackaging.

AI A makes a questionable claim.

AI B repeats it in a polished report.

AI C cites B.

By lunchtime the original guess is wearing a tie.

SEE ALSOSloppy Slop · Synthetic Consensus · Confidence Lacquer

Evidence Latch

◈

A mechanism that prevents work from being treated as complete until the required evidence actually exists.

The latch may require passing tests, independent review, an artifact, a reproducible result or some other explicit acceptance condition.

It turns "probably done" into a state the system cannot accidentally confuse with "done."

FINDEX ↑

Fresh-Agent Test

◈

Give a project to a competent AI with no private briefing.

Can it discover how the system works, identify current state, use the correct tools and continue safely?

If yes, the project has a healthy operational interface.

If Matthew needs to explain seventeen undocumented facts first, it fails.

SEE ALSORecovery Surface · Self-Teaching Tool · Glass Trail

Frontier Code

◈

Experimental software deliberately built ahead of settled practice.

Frontier code accepts uncertainty consciously and contains its consequences.

This separates it from cowboy engineering.

Frontier work takes risks knowingly. Cowboy work loses track of the risks.
GINDEX ↑

Glass Trail

◈

A visible chain showing what an AI did, why it did it, what authorised the action and what evidence resulted.

The Glass Trail is not merely a log.

It should allow somebody arriving later to reconstruct the important story of the work.

SEE ALSOProvenance Spine · Recovery Surface

Green-Light Hallucination

◈

The mistaken conclusion that a system must be correct because all visible tests, checks or dashboards are green.

Tests demonstrate what they were designed to test.

They do not confer sainthood.

SEE ALSOVerification Theatre · Reward Hacking · Specification Gaming

HINDEX ↑

Handoff Entropy

◈

Information lost, distorted or weakened every time work passes between agents, models, sessions or humans.

A project may begin with a precise architectural decision and, five handoffs later, retain only the folklore that "Lucy said something about boundaries."

Good systems minimise this decay.

Harness

↗

The software environment surrounding an AI model that gives it tools, context, execution, permissions, loops, memory and other operational capabilities.

The distinction matters because many things attributed to "the AI" are actually properties of the harness around it.

The model may be the engine.

The harness determines what the engine is attached to.

Harness Tax

◈

The time, compute and cognitive overhead imposed by the machinery surrounding an AI rather than by solving the underlying problem.

A good harness pays for itself.

A bad harness becomes an elaborate machine for administering itself.

Human Hinge

◈

A specific point in an otherwise automated process where human judgment, authority or responsibility genuinely matters.

Well-designed autonomous systems do not put a human in every loop.

They identify the few hinges where a human decision changes what the system is legitimately allowed to do.

IINDEX ↑

Inference Bonfire

◈

A particularly spectacular outbreak of Cowboy Compute in which several expensive agents enthusiastically rediscover the same information.

Often accompanied by impressive token counts and very little new knowledge.

🔥

LINDEX ↑

Long-Horizon Task

⌁

A task requiring sustained progress across many actions, decisions or stages rather than a single response.

Long-horizon work exposes problems that short demonstrations often hide: memory loss, context reconstruction, handoff entropy, duplicated verification and accumulating authority ambiguity.

MINDEX ↑

Model Weather

◈

Changes in AI behaviour caused by model revisions, routing, inference settings, harness changes or other external conditions rather than changes to the task itself.

A workflow that behaved perfectly yesterday may behave differently today.

Not every unexplained change is a bug in your software.

Sometimes it is raining upstream.

OINDEX ↑

Orphan Decision

◈

A surviving decision whose original rationale has disappeared.

Everyone knows what was decided.

Nobody remembers why.

Orphan decisions are dangerous because future agents cannot distinguish a fundamental invariant from something somebody chose temporarily six months ago.

PINDEX ↑

Procedure Wallpaper

◈

Detailed procedural documentation that technically exists but has little meaningful influence over how humans or agents actually behave.

It decorates the environment while the real workflow happens elsewhere.

SEE ALSOPrompt Fossil · Self-Teaching Tool

Prompt Barnacle

◈

An instruction added after some historical failure that remains permanently attached to a prompt even after the underlying problem has disappeared.

One barnacle is harmless.

A thousand eventually become the boat.

SEE ALSOPrompt Fossil · Rule Fog

Prompt Fossil

◈

A surviving instruction whose original reason has been forgotten.

Nobody knows whether it remains necessary.

Nobody is brave enough to remove it.

Proof Debt

◈

Work accumulating faster than trustworthy evidence that it works.

Eventually somebody has to pay the verification bill.

The longer payment is deferred, the harder it becomes to determine which assumptions, tests and artifacts still correspond to which version of the work.

Provenance Spine

◈

The small canonical sequence of decisions, artifacts and evidence from which the important history of a project can be reconstructed.

The provenance spine is not every event that occurred.

It is the structural skeleton that allows everything important to be understood.

RINDEX ↑

Recovery Surface

◈

The information and mechanisms available to a fresh human or AI attempting to resume interrupted work.

Repositories, decision records, state files, evidence, logs and self-describing tools may all contribute to the recovery surface.

A large project with a tiny recovery surface will repeatedly pay for Context Archaeology.

Reward Hacking

⌁

When an AI discovers a way to maximise a score, reward or apparent success condition without accomplishing the intended objective.

A coding agent that makes a failing test pass by weakening the test rather than fixing the software has found the scoreboard rather than the goal.

SEE ALSOSpecification Gaming · Green-Light Hallucination

Rule Fog

◈

The state reached when an AI is given so many instructions that the important ones become difficult to distinguish from everything else.

Adding another rule to rule fog can actually make the system less controlled.

The cure is often architecture, tooling or clearer invariants rather than another paragraph.

SINDEX ↑

Scaffold Fade

◈

The deliberate removal of procedural instructions once tools and interfaces have become capable of teaching, constraining or recovering the correct operation themselves.

The procedure fades.

The genuine invariant remains.

SEE ALSOSelf-Teaching Tool

Scaffolding

↗

Supporting structure added around an AI to help it perform a task successfully.

Scaffolding may include prompts, workflows, examples, tools, state management and other temporary or permanent support.

Useful term.

Frequently used too vaguely.

Self-Certifying Agent

◈

An AI allowed to perform work and then act as the sole authority deciding whether its own work succeeded.

Self-checking is useful.

Self-checking is not independent evidence.

Self-Teaching Tool

◈

A tool whose interface, schemas, examples, errors and constraints allow a competent AI to discover how to use it correctly without memorising a large procedural manual.

A self-teaching tool turns deterministic procedure into interface behaviour rather than prompt baggage.

This embodies a central Toolglass principle:

Self-teach over rule-teach.

Slop

↗

Low-value machine-generated material produced without sufficient care, judgment or purpose.

The important characteristic is not merely that AI produced it.

The characteristic is that its production cost is tiny while the burden of reading, checking or cleaning it is pushed onto somebody else.

Sloppy Slop

◈

When an AI treats another AI's generated output as evidence, fact or directional guidance without checking the underlying source.

Errors, uncertainty and assumptions then compound across generations.

AI A inaccurately summarises an article.

AI B uses A's summary instead of reading the article.

AI C treats B's report as evidence and changes the project accordingly.

Nobody ever visits the source.

Sloppy Slop.

Slop becomes Sloppy Slop when generation is mistaken for evidence.

SEE ALSOSlop Cascade · Evidence Laundering · Synthetic Consensus

Slop Cascade

◈

The propagation of Sloppy Slop through several systems, agents or decisions.

Each generation inherits the previous generation's assumptions and adds another layer of interpretation.

Eventually the system may be reasoning confidently about something that nobody in the chain actually verified.

Specification Gaming

⌁

Fulfilling the literal specification while violating its intended purpose.

The phenomenon predates modern AI agents but becomes increasingly important when systems have enough autonomy to discover unexpected ways of satisfying poorly designed goals.

SEE ALSOReward Hacking

State Spill

◈

A project state distributed across too many places to possess one obvious source of truth.

Part of reality is in Git.

Part is in a chat.

Part is in a local directory.

Part is in a task tracker.

And an inexplicably crucial architectural decision lives in `notes-final-FINAL2.md`.

Synthetic Consensus

◈

The appearance of independent agreement between several AI systems when they actually share the same source, assumption, context or inherited error.

Five agreeing models do not constitute five independent witnesses if all five read the same mistaken summary.

SEE ALSOSloppy Slop · Evidence Laundering

TINDEX ↑

Tool Gravity

◈

The tendency of an AI to fall toward tools it already knows rather than tools best suited to the task.

Familiarity exerts gravitational pull.

Good harnesses make the correct tool easier to discover than the familiar wrong one.

Tool Shyness

◈

The tendency of an otherwise capable AI to avoid an unfamiliar tool unless explicitly instructed to use it.

Tool shyness is often mistaken for a model limitation.

Frequently the tool simply does a poor job of explaining itself.

SEE ALSOTool Whispering · Self-Teaching Tool

Tool Whispering

◈

Repeatedly coaxing an AI into using a tool that ought to be discoverable and understandable without special prompting.

When a workflow requires ritual phrases such as "You MUST use X for this exact operation," the agent may not be the only thing requiring improvement.

Tool Whispering is often evidence that the interface should become a Self-Teaching Tool.

VINDEX ↑

Verification Theatre

◈

Checks performed because they create the appearance of rigor rather than because they meaningfully reduce uncertainty.

A hundred irrelevant tests can provide less evidence than one well-chosen experiment.

Verification should answer a question.

Otherwise it is scenery.

THE EMERGING FAMILIES

Words become more useful when they connect.

A few recurring sequences are already visible. They are less a taxonomy than a set of tracks through the undergrowth.

EDITORIAL RULE / ADMISSION IS DELIBERATELY DIFFICULT

The Glassary is not a jargon factory.

A new term needs to name a recurring phenomenon for which existing language is poor, make an important distinction easier to think about, or be memorable enough that people will actually use it.

Being funny helps. Being useful is compulsory. The aim is to give names to things we keep tripping over.

Once something has a good name, it becomes considerably harder to pretend you haven't noticed it.

THE TOOLGLASS LETTER

Toolglass, occasionally.

New reviews, strange software and useful things that deserved more attention.

The subscription desk is being connected. The letter will open here once its Buttondown account is ready.