The Context Wall
Diagnose context failures, organize durable technical knowledge and add QMD retrieval using documented commands, explicit access boundaries and observable maintenance.
A project context file can be a useful starting point for an AI coding workflow. It holds the decisions, conventions and domain facts that generic model knowledge cannot supply. As that material grows, the question becomes how to make the right evidence available for each task.
There is no universal line count at which the file stops working. Lines vary in length, clients load different context, and task difficulty matters. Diagnose missing or misused evidence before deciding whether to reorganize documents, add retrieval or change the workflow.
Separate storage from active context
The durable knowledge base is what the project preserves. Active context is the subset supplied to a model request, along with instructions, conversation and tool output. A large repository does not need to be loaded into every request.
A short context file offers transparency and little infrastructure. Its contents still need to be correct, relevant and actually read. Information being present does not guarantee that a model will use it accurately.
As the archive grows, a short overview can guide navigation while detailed documents hold the evidence. Long-context reading and retrieval are compatible techniques; neither is the correct answer for every task.
Diagnose the problem with representative questions
Look for recurring failure cases: an outdated decision used as current, a relevant rule omitted, or an answer that cites a source without supporting the claim. Record the expected evidence and inspect whether the agent retrieved and read it.
Check actual context usage with the client’s supported diagnostics. Avoid converting line count to tokens with a universal multiplier or adding a fixed assumed conversation overhead. The active prompt can differ substantially between sessions.
Use several representative questions, including exact terms, synonyms and facts in different locations. A single wrong answer does not prove context exhaustion; it may reflect a retrieval failure, stale source, ambiguous question or reasoning error.
Preserve meaning when splitting files
Use topic and record boundaries rather than an arbitrary maximum line count. Keep essential definitions close to the material that depends on them, and maintain stable links between decisions, evidence and related concepts.
project/
CLAUDE.md
docs/
overview.md
architecture/
decisions/
interfaces/
investigations/This layout is illustrative. Configure the client to read the relevant overview or open it explicitly; its filename alone does not guarantee automatic loading. Keep authoritative shared documentation in the location approved by the project.
Add retrieval when the evidence supports it
QMD’s upstream documentation describes keyword search, vector search and hybrid retrieval with reranking. These provide ways to discover candidate documents. A retrieved passage remains evidence to inspect, not a certified answer.
Keyword search is useful for exact identifiers and terms. Semantic retrieval can find related wording. Combining and reranking candidates can improve some workloads but introduces its own latency and failure modes. Evaluate those tradeoffs on the corpus being used.
Precise latency claims need hardware, corpus size, configuration and warm-versus-cold conditions. This article does not promise millisecond queries or a fixed startup delay. Nor does it require an adoption statistic to establish whether the tool is useful.
A documented installation path
Checked against the upstream package documentation on 20 September 2026 (package version 2.8.3), the intended package is @tobilu/qmd. Verify the current release and supported runtime before installation; pin an approved version where reproducibility is required.
npm install -g @tobilu/qmd
qmd collection add ./docs --name project-docs
qmd update && qmd embed
qmd search "retry policy" -c project-docs
qmd query "Why was the retry policy chosen?" -c project-docsThis registers a collection and performs maintenance explicitly. It does not use the unrelated mempalace package, qmd init, qmd index or an invented project .qmdrc file. Inspect any collection update commands before running them, because maintenance configuration can include executable hooks.
QMD’s current source documents optional AST-aware chunking for supported code files using --chunk-strategy auto; Markdown uses its documented text chunking. Treat model choices and chunking behavior as versioned implementation details, not permanent properties.
Connect the client in the correct scope
For Claude Code, a project MCP configuration belongs in .mcp.json. The official MCP guide distinguishes project, local and user scopes and their trust behavior. Do not assume a server block can be placed interchangeably in every settings file.
For an installed qmd executable on the client’s PATH, the following is a project configuration fragment. Merge it into the existing object; do not overwrite unrelated servers. Review and approve project-provided servers according to the client’s trust controls.
{
"mcpServers": {
"qmd": {
"command": "qmd",
"args": ["mcp"]
}
}
}The upstream tool surface includes query, get, multi_get and status. Confirm the discovered tools in the running client, then test an authorized query and retrieval of its source document. Successful registration alone does not prove that the index contains the expected material.
Choose a service lifecycle deliberately
QMD supports a persistent HTTP mode. That can avoid repeatedly starting a process, but it also creates a service to operate. Bind it only where intended, constrain who can connect and use a supervisor if the workflow needs continuous availability.
qmd mcp --httpA collection filter is not access control when the caller can query the entire shared service. Local retrieval also does not mean a cloud agent keeps returned passages local. Trace the complete data flow and enforce the required boundary before combining corpora.
Treat freshness as a checked dependency
When files change, the search index and embeddings may need maintenance. Use a single coordinated worker for a shared index. Commit hooks should enqueue or coalesce requests under the same locking policy rather than launching overlapping writers.
Keep pending requests durable. Require a successful update before embedding, and mark a request complete only after the relevant checks succeed. If a new request arrives while work runs, retain it for another pass. Preserve failure evidence for diagnosis.
A safe maintenance design also coordinates upgrades with index writes. Back up the required state and preserve a recovery path before changing runtime dependencies. Installation success must precede a dependent restart; a failed health check must return a failure status.
Check the subsystem that failed
A diagnostic command should distinguish executable health, service connectivity, index freshness and retrieval quality. A successful qmd status cannot by itself clear a daemon or embedding failure.
Use the client’s actual MCP connection check and a known-source query for the service path. Verify the running version after a restart. Clear only the recorded failure whose recovery was tested, and keep unresolved errors visible.
These checks improve detection. They do not make silent failure impossible, and a test probe can itself have blind spots. Review the monitoring when a new failure mode appears.
Guide retrieval without disabling useful investigation
Instructions should tell the agent when to consult project knowledge and when exact search is appropriate. A broad regular expression over shell command text cannot reliably distinguish every search from quoted text or every path form.
If a policy requires a search guard, implement it against the supported tool input and test relevant command forms, quoting and path cases. Document any fail-open behavior. Retain an authorized exact-search fallback for known strings and diagnostics.
A denied grep call is not proof that the agent subsequently used retrieval. If that evidence matters, record the actual retrieval request and the sources consulted, with suitable data minimization.
Compare the resulting workflow
Run the same representative questions through the baseline and the revised setup. Assess source recall, answer support, freshness, latency and operational effort. Investigate misses rather than assuming that another index or larger context always improves quality.
Keep the simple file when it meets the need. Add retrieval when it provides measurable value, and preserve source portability so that the implementation can change later. Larger context windows may alter the tradeoff; no fixed forecast is needed to choose today’s system.
The useful outcome is an agent that can recover the right evidence and a maintainer who can tell when that process failed. Durable documents, explicit scope and observable maintenance matter more than a universal threshold.