Dev Content Pipeline: Making Repository Architecture Visible with AI
A developer tool that combines repository analysis and AI to prepare architecture diagrams and engineering articles, supported by source excerpts and reviewed by the author.
Published:
Updated:

When I reopen a project after some time away, I often remember what it does before I remember how its parts fit together. Reconstructing those relationships takes another pass through the code—and explaining them to someone else takes another kind of work. I am developing Dev Content Pipeline to help connect that investigation to a readable architecture explanation, with Mermaid diagrams and source excerpts I can inspect.
Returning to a repository I need to understand again
Dev Content Pipeline is a personal engineering tool that I actively develop and already use locally to prepare Engineering Lab content. The most useful output for me so far is the combination of a diagram showing relationships and relevant code showing what supports them. Together, they help me recover context and communicate it.
At the time I wrote the brief for this case study, all my published Engineering Lab articles had been prepared using the tool. That is my account of how I use it. My intention is to prepare this article through the same workflow: use the pipeline that helps me explain other projects to explain the pipeline itself.
AI-assisted prose is part of that process, but I still choose the material, review the claims, and edit the result. The practical goal is to arrive at an explanation whose important relationships I can check against the repository.
A brief gives the analysis a purpose
I start with a Markdown brief containing repository paths, remembered context, and the direction I want the article to take. That context gives the investigation a purpose. A recollection can suggest where to look without becoming proof of how the system works.
The CLI separates the workflow into three commands, connected by saved artifacts:
- Analyze receives the brief path and writes a research directory.
- Compose reads that directory, creates and validates an editorial plan, then requests prose.
- Generate reads the composed material and builds a separate publication package.
The V3 execution documentation describes analysis as deterministic file discovery and topic matching, accompanied by bounded source snapshots. It records four research artifacts covering the original brief, observations, source text, and configuration. The analyzer implementation is outside the supplied source bodies, so those collection details remain documentation-based. Keyword matches provide leads for reading related code; they do not establish a dependency graph.
The composition code does show that the complete original brief reaches planning alongside requested diagrams, findings, research text, and snapshots. The first model request proposes the story and its supporting relationships. Local validation checks that plan, and the pipeline saves it before requesting the article body.
This is a locally run workflow with an external composition boundary. The adapter launches a Codex executable in a temporary directory, requests read-only sandboxing, disables shell tools, and supplies a JSON response schema. It sends the prompt through standard input and reads a response file. As the documentation explains, the brief and filtered source snapshots go to the authenticated model service. Local analysis and rendering therefore do not make the process offline.
These boundaries also explain the overall architecture diagram: some arrows represent calls, while others represent files written by one command and read by a later invocation.
Following one relationship from code to diagram
A useful diagram relationship needs a precise verb. In this repository, a concrete example is generation calls the shared Mermaid renderer on its staged package.
The supporting code spans two responsibilities. generateContent writes the package into a staging directory and calls the imported renderDirectory(draft) function. That renderer recursively discovers .mmd files and invokes Mermaid CLI’s imported run function to create adjacent SVG files. Generation uses the shared function directly; the standalone diagrams command is another entry point to that renderer.
This makes a small but meaningful architectural distinction visible: rendering happens within generation, while the Mermaid sources arrive from composition through saved artifacts.
The diagram specification represents such a relationship as two declared components and a labeled edge. The edge carries its evidence basis and references to captured source ranges. Boundaries carry evidence too. Validation checks that references fall within the available snapshots, that edge endpoints exist, and that nodes refer to declared boundaries.
After composition, renderArchitecture turns the specification into Mermaid text. It constructs nodes and boundary subgraphs, escapes labels, and assigns colors by component kind. An edge labeled as an inference gets a dashed arrow. Other supported bases use solid arrows, including documentation and configuration, so line style alone cannot establish runtime behavior.
Before generation proceeds, the publication reader checks that the saved specifications equal the plan and that each .mmd file equals the deterministic rendering of its specification. Generation then follows the staged rendering path described above. This establishes how a proposed relationship reaches an SVG; it does not establish that this article’s SVG has already been rendered.
The distinction matters during review. Validation can confirm that an arrow points to existing components and cites available lines. I still need to read those lines and decide whether they support the arrow’s meaning.
const files = renderEditorialPackage(publication, article, card);
await mkdir(dirname(output), { recursive: true });
// Reserve the destination before browser work; a failed conversion leaves no partial package.
await mkdir(output);
let staging: string | undefined;
try {
staging = await mkdtemp(resolve(dirname(output), ".content-"));
const draft = resolve(staging, "draft");
await writePackage(draft, files);
const svgCount = await renderDirectory(draft);
// Rename onto our empty reservation. Other drafts are never replaced.
if (!(await lstat(output)).isDirectory())
throw new Error(`Invalid output: ${output}`);dev-content-pipeline/src/generate-content.ts:17–29
export async function renderDirectory(directory: string): Promise<number> {
let count = 0;
const entries = await readdir(directory, { withFileTypes: true });
for (const entry of entries.sort((a, b) => a.name.localeCompare(b.name))) {
const input = join(directory, entry.name);
if (entry.isDirectory()) count += await renderDirectory(input);
else if (entry.isFile() && entry.name.endsWith(".mmd")) {
const output = `${input.slice(0, -4)}.svg` as const;
const existing = await lstat(output).catch(
(error: NodeJS.ErrnoException) => {
if (error.code !== "ENOENT") throw error;
return undefined;
},
);
if (existing && !existing.isFile())
throw new Error(`Output is not a regular file: ${output}`);
await run(input, output, { parseMMDOptions: { backgroundColor: "transparent" } });
count += 1;
}
}
return count;dev-content-pipeline/src/render-diagrams.ts:6–26
for (const edge of diagram.edges)
lines.push(
` n_${edge.from.replaceAll("-", "_")} ${edge.basis === "inference" ? "-." : "--"} "${label(edge.label)}" ${edge.basis === "inference" ? ".->" : "-->"} n_${edge.to.replaceAll("-", "_")}`,
);dev-content-pipeline/src/editorial-diagrams.ts:29–32
Keeping revisions and publication assets inspectable
Saved intermediate results make the workflow easier to inspect and resume. With --plan-only, composition stops after persisting the validated plan. With --from-plan, it reloads and validates that plan against the research package before drafting. I can review the proposed story and relationships, or retry drafting, without requesting another plan.
Freshness checks keep these artifacts tied to their inputs. The editorial reader compares the current brief with its recorded hash and computes an input hash from the four research-file strings. Source verification rereads each full file, compares its hash, reconstructs the captured line range, and checks that the snapshot text matches.
Composition verifies all supplied snapshots. Generation verifies the sources referenced by claims, selected excerpts, diagram edges, and boundaries. Because source hashes cover full files, a change outside a displayed excerpt can still make its evidence stale. These checks detect changed inputs; they do not assess the quality of the explanation.
Selected excerpts are extracted directly from snapshots during package assembly. Their repository paths and line ranges accompany them, while planned diagram positions become SVG references. The package includes MDX, Mermaid sources, rendered SVGs, HTML excerpt cards, LinkedIn text, the plan, a manifest, and review notes.
Rendering also has a clear completion boundary. Generation reserves a new destination, writes and renders in staging, then renames the completed draft into place. Its failure path removes the reservation, and cleanup removes staging. A rendering failure therefore does not intentionally leave a partial publication package at the destination.
I publish only a selection of this material. I review, edit, and shorten it for a broader Engineering Lab audience. Fuller architecture explanations and source snippets can still help developers or architects examine the implementation closely. A shorter article reflects my editorial choice about how much to present.
export async function verifySources(
directory: string,
input: EditorialInput,
sources: EvidenceSource[] = input.sources,
) {
for (const source of sources) {
const root = input.brief.repositories[source.repository];
if (!root) throw new Error(`Unknown repository: ${source.repository}`);
const full = await readSelectedFile(resolve(directory, root), source.path);
if (hash(full) !== source.sha256)
throw new Error(
`Source changed since analysis: ${source.repository}/${source.path}. Rerun analysis and compose.`,
);
const snapshot = full
.replace(/\r\n/g, "\n")
.replace(/\n$/, "")
.split("\n")
.slice(source.startLine - 1, source.endLine)
.join("\n");
if (snapshot !== source.text)
throw new Error(`Evidence text does not match source: ${source.id}`);
}
}dev-content-pipeline/src/editorial-input.ts:79–101
What still needs editorial judgment
My experience is that generated prose can be more expansive than I want. Editorial focus and concision remain areas for improvement, even when the diagrams and supporting excerpts are useful.
Source coverage also limits the explanation. The V3 documentation describes snapshots capped at 240 lines and 16,000 characters, with full-file hashes and recorded line counts that expose truncation. A valid reference can only reach captured text. Missing evidence cannot establish that a relationship does not exist, and structurally valid evidence can still be misinterpreted by the model.
The current composition path already requests LinkedIn prose, saves linkedin-draft.md, and carries that text into the generated package as linkedin.md. I would like to improve the assistance around that draft: identifying the main point worth sharing, writing a concise description, choosing relevant links, and suggesting visuals from the prepared article material. A diagram or source excerpt could give a short post something concrete to discuss.
Those are development ideas, not completed capabilities or a committed delivery plan. I would still decide what to say, check claims and links, choose visuals, and approve the final material before posting.
That same responsibility remains central to the article workflow. The pipeline gives me an inspectable chain from a brief through source evidence to an architectural explanation. Its value depends on my ability to follow that chain and judge whether the explanation holds.