# What search engines and AI assistants read on a page

> The invisible half of a page decides whether it is trusted. Here is what we put there and why.

Published 2026-08-18 by Searchlode.

A visitor sees a heading, an answer and a table. A crawler sees much more, and the difference decides whether the page is treated as reference material or as noise.

## robots.txt and the sitemap

The first two things a crawler fetches. The sitemap lists every published page with the date it last changed; pages that have been held back by the quality check are never listed, so nothing thin is ever offered for indexing.

## Structured data

Every page declares, in machine-readable form, who wrote it, who reviewed it and with what credential, where the numbers come from, what the page is about, and which sentence is the answer. Google reads this for its result features; assistants read it to decide whether to cite the page.

## The Markdown twin

Every page also exists as plain text at the same address with .md on the end, and the subdomain publishes an index of them. That is the format AI assistants prefer to ingest, and publishing it is the cheapest way to be quotable.

## The status codes

A live page answers 200. A discontinued service's page answers 410 (gone), so search engines drop it cleanly instead of finding a broken link. Housekeeping like this is most of what “technical SEO” means in practice.

Canonical: https://www.searchlode.com/insights/what-search-engines-and-ai-assistants-read
