RAMYA LAKHANI

All log entries

What I was doing. Rendering answers from a language model into the page, with the quoted source excerpts they were built from.

What I noticed. The usual advice is to escape untrusted text before rendering it. On this corpus that is the wrong instruction, because the content is attack strings. An excerpt from a writeup legitimately contains <script>, ../../, ' OR 1=1, and a Authorization: Bearer header. Escaping mangles the thing the reader came to read; rendering it as markup is the exact bug class the page is about.

The way out is that neither is necessary. There is a third option that is both safe and lossless:

node.textContent = value          // never innerHTML, never insertAdjacentHTML

textContent does not parse. A <script> in an excerpt stays five visible characters — correct on screen and inert. The whole island is written this way, and the rule is stated at the top of the file so the next person editing it knows it is load-bearing rather than stylistic.

The second half was citations. The model returns a URL per claim, and a citation that points off-site is indistinguishable to a reader from one that does not. So a URL is validated twice, in two different processes: the server anchors it to the site origin, and the client independently re-parses it and requires the path to match a real archive route before it becomes an href.

Why it matters for shipped code. Any feature that renders model output has this shape, and retrieval-augmented ones have it twice — the generated text and the retrieved source are both untrusted, and the retrieved source is the one people forget. Two habits transfer: decide once, at the boundary, that this data never becomes markup, and re-validate anything that becomes a link in the process that renders it, not only in the one that produced it.

Open question. Streaming. Appending tokens as they arrive means the safe-render rule has to hold on every partial write, and I have not yet convinced myself the obvious implementations do.