Field test · NJ School Data
A real AI visibility audit, end to end
This is one observed test, not a polished success story. It shows the full chain from an audience need to a change in the source product—and preserves the null result that came next.
1 · Choose a question for a reason
Start with the decision a reader is trying to make
NJ School Data had 201 questions gathered from its audience work. We selected 12 that together covered every audience role, decision function, and content job in that set. Twelve was the smallest complete baseline panel—not a product limit or an arbitrary quota.
“What happens to other kids here with needs like my child’s—is there a pattern I should know about?”
2 · Separate source problems from tool problems
Audit the diagnosis before prescribing work
The first pass said the matched data page lacked structured data. The live source already had appropriate Dataset markup; our extractor had missed it. We repaired that diagnosis instead of turning our own defect into work for the publisher. The page did still lack a direct parent-facing answer and exact freshness metadata.
3 · Change the real publishing product
Ship the smallest source-backed intervention
We added a bounded explanation of what the district-level figures can and cannot show, plus exact Dataset and freshness metadata, to the live discipline-and-safety page.
See the changed page on NJ School Data →
| Measure | Before | After |
|---|---|---|
| Candidate relevance | 0.70 | 0.90 |
| Page diagnosis | 0.536 · needs work | 0.893 · good |
| Direct answer | Fail | Pass |
| Dataset schema | Pass | Pass |
| Freshness | Fail | Pass |
4 · Keep the null result
The page improved. Ambient discovery did not.
The same 12 questions were retested once across three live answer surfaces. NJ School Data had 0 citations in 36 observations before the intervention and 0 in 36 after it. For this question, all three systems treated “here” and “needs like my child’s” as missing context. Better page readiness did not create immediate visibility.
5 · Test the next explanation separately
Explicit district context produced a directional result
We preserved the broad question as the ambient-discovery measure and added a separate South Orange-Maplewood version. On that grounded prompt, one of three surfaces cited the exact district page; two did not. That is evidence that the page can be retrieved for specific intent. It is not evidence of repeatable citation lift.
| Surface | Model or interface | Observed result |
|---|---|---|
| Codex through Somm | gpt-5.6-sol | Cited the district page first |
| Claude through Somm | claude-sonnet-4-6 | Did not cite it |
| Google AI Mode | Rendered answer surface | Did not cite it |
What the test established
A useful audit leaves an inspectable chain
This test established that we can derive questions from audience needs, distinguish an internal extraction failure from a real source deficiency, implement a source change, and measure page readiness and retrieval separately. Each observation retained its question, surface, model or interface, date, citations, matched source, and limitations. MiniMax M3 was used for source matching, not answer evaluation; Nomic supplied local embeddings.
It did not establish repeatable visibility lift, referrals, customer value, demand, payment, or revenue. Those are separate tests. The next commercial question is whether a publisher values this evidence-to-implementation loop enough to buy it.
Pilot engagements
Bring us a publishing problem, not a keyword quota
We are looking for evidence-rich publishers and information products that want to understand—and improve—how AI answers use their work.
Talk with Lyra Forge