Key takeaways
- 12 French SEO sector sites crawled end to end on 27 August 2026: 8,246 pages opened one by one, 95,442 content links recorded, 9 sites in the comparable panel
- The homepage reaches a median of only 58.6% of sitemap pages through content links present in the served HTML. The spread runs from 3.5% to 95.1%
- Shallow depth plus high concentration do not travel together: Spearman's rho is -0.289 across 9 sites, +0.025 across 8. The expected siloing signature does not appear in this panel
- What sits out of reach is always the same thing: other languages, the paginated blog, the glossary. The link exists, it just lives in the nav
- Orphan and unreachable are two different failures: seo.fr has 0.4% orphan pages and 41.4% of pages no content-link path connects to the homepage
Semantic siloing is a French doctrine. It promises that a strict tree, parent pages that distribute and child pages that point back, concentrates relevance and lifts a target page. Sold for over a decade by nearly every agency in the country, it is still the default architecture reflex in French SEO.
The 2026 question is not whether it was ever true. It is what survives when a growing share of journeys starts inside ChatGPT, Perplexity or an AI Overview, and those engines do not read a site the way Googlebot does.
We chose to measure rather than argue. On 27 August 2026 we opened 8,246 pages one at a time across twelve French SEO sector sites, and recorded 95,442 content links. Same method, same day, same definition for everyone.
The one-line result: the expected siloing signature does not appear in this panel, and the real architecture gap sits elsewhere.
Semantic siloing, in one sentence
Semantic siloing is a content tree where each page covers one narrow sub-topic and links only to its immediate neighbours in the hierarchy, so that link equity and topical relevance concentrate on a target page.
Two properties follow mechanically, and both are measurable:
- shallow, regular depth from the homepage, because the hierarchy is short and complete;
- heavy concentration of inbound links on a handful of head pages, because that is the whole point.
Which gives a question a crawl can answer: do those two properties show up together on real sites?
One clarification that is not cosmetic: a crawl says nothing about intent. No site in this measurement is labelled here as siloed, cocooned or flat. We publish profiles, not labels.
What we measured, and what we did not
Fifteen domains targeted, twelve measured. For each: read the sitemap, sitemap indexes included, then open every URL, no sampling. A link counts as a content link if it sits in the page's content zone, outside nav, header and footer, and points at another sitemap page returning 200. The HTML is taken exactly as served, no JavaScript is executed.
A site only enters the comparable panel if its content zone was detected through the main tag on at least 95% of its pages, fewer than 5% of its URLs fail, and none stays stuck on 429 or 403. Nine sites pass, three are published with an explicit caveat and outside every benchmark, three are excluded: resoneo.com returns 403 down to its robots.txt, open-linking.com and 1ere-position.fr both redirect to noiise.com, already measured.
Two domains in the panel are not neutral for us, so let us say it here: vydera.com is our own site and digidop.fr is a Vydera client. Both run through the same script, the same rules and the same thresholds as the other seven, and their figures are published as they came, including the awkward ones.
What the measurement does not say, and what we therefore write nowhere:
- Nothing about AI engine retrieval. No citation, no passage retrieval was measured here. This dataset cannot support the claim that one architecture yields more citations than another.
- Nothing about the final DOM. A page out of reach in this measurement may well be linked from a JavaScript-rendered menu. It is out of reach of the served HTML, which happens to be the perimeter most robots read.
- Nothing about traffic, rankings or authority for these domains. No performance is compared.
- Nothing about time. One snapshot, one day.
And one methodological caveat heavier than all the others: the definition of the content zone moves everything. On digimood.com, same site, same day, depending on whether content is detected through the article tag or by subtracting nav, header and footer, the figures shift from 1.91 to 11.14 links per page, from 39.5% to 27.6% orphan pages and from 55.0% to 88.6% concentration. Direct consequence: two architecture figures published by two different tools are not comparable.
What sits out of reach of the homepage is never a surprise
First figure, and the loudest. Across the nine comparable sites, the homepage reaches a median of only 58.6% of sitemap pages through content links present in the served HTML. Put another way, 41.4% of the site stays outside the graph for a robot that runs no JavaScript. The spread between sites is enormous: from 3.5% on primelis.com to 95.1% on agence-slashr.fr, same day, same method.
The useful part is not the median, it is the nature of what falls out of the graph. It barely varies:
- seo.fr: 206 unreachable pages, of which 204 sit under the
/enprefix. The English version exists and is linked, but from the language switcher. - keyweo.com: 739 unreachable pages out of 1,589, dominated by
/de(241),/it(212) and/nl(165). Same cause, across five languages. - natural-net.fr: 344 unreachable, of which 336 under
/blog-agence-web. A paginated blog index that never unfolds in HTML. - vydera.com: 80 unreachable, of which 74 glossary pages, because the only page listing them receives no content link itself.
Three mechanisms, always the same three: the language switcher, the paginated index, and the hub that only lives in the nav. None of them is an architecture doctrine failing. They are links that exist but sit where a robot does not look.
That is the most profitable operational conclusion in the whole dataset: before redrawing a tree, check that the served HTML carries the links. The method side is covered in our article on internal linking.
The siloing signature does not appear in this panel
Siloing should produce two effects together: shallow depth and heavy concentration. So we checked whether the two move together.
They do not. Spearman's rho between median depth and concentration is -0.289 across the 9 sites, and +0.025 across the 8 once the degenerate case described below is removed. A coefficient that slides from weakly negative to flat zero when one site leaves is noise, not a relationship.
Read that result for what it is. n equals 9. This is an observation on nine sites, not a statistical proof, and none of these coefficients carries test value. What we can write: the nine measured sites do not show shallow depth paired with heavy concentration. What we cannot write: siloing does not work.
Median depths run from 1 to 7 clicks from the homepage, with a median of 3. Maximum depths run from 1 to 17. Concentration spreads from 31.2% to 94.9%, median 55.2%. With template links stripped, the panel median barely moves, to 56.4%. In this panel these are two independent dimensions, not two faces of one architecture.
One last caution, and a serious one: platform probably weighs more than SEO philosophy on these numbers. Five WordPress, two Webflow, one Next.js, one unidentified platform. There is no way to separate an architecture choice from a theme.
Orphan and unreachable are not the same thing
Most audits report a single figure, orphan pages. A second one is missing, and the measurement shows it plainly.
An orphan page receives no content link at all. An unreachable page receives some, but no content-link path connects it to the homepage. The two barely overlap.
seo.fr shows 0.4% orphan pages and 41.4% unreachable pages. A standard orphan audit would hand it a clean bill. Its 204 English pages do receive links, they link to each other, but the whole set forms a continent detached from the homepage.
Across the panel, the orphan share runs from 0.3% (digidop.fr) to 76.9% (primelis.com), median 6%. Two sites clear 29%: eskimoz.fr at 29.4% and natural-net.fr at 46.5%. Track both indicators, not one.
The edge case: a homepage that reaches six pages
primelis.com earns its own paragraph, because it sets five of the panel's nine benchmarks and shows what a site with no editorial linking in its HTML looks like.
170 sitemap pages, 514 content links in total, that is 3.02 links per page. Of those 170 pages, 87 serve exactly 5 content links and 59 serve none at all. Five offer pages absorb nearly all of them: /paid-media, /retail, /organic-growth and /data each receive 89 inbound links, /creative-studio 83.
Consequence: maximum depth of 1 click. The homepage reaches six pages, itself included, and the graph stops there. 164 of 170 pages stay out of reach, 76.9% of the site is orphan, concentration climbs to 94.9% and Gini to 0.948.
Because such a site distorts every benchmark, the dataset publishes them twice, under a mechanical rule rather than a judgement call: any site whose homepage reaches under 10% of its sitemap is removed. Without primelis.com, median concentration hardly moves, from 55.2% to 55.0%, but the maximum drops from 94.9% to 69.9% and the minimum reachable share climbs from 3.5% to 50.2%.
Why an LLM does not read a silo the way Googlebot does
What follows is reasoning about mechanism, not a measurement. Nothing in this dataset covers what an AI engine fetches or cites. We write it because the question is fair, not because we measured it.
A classic search engine walks a graph. It follows links, propagates authority from page to page, and architecture acts on that propagation: orchestrating it is exactly what siloing is for.
A generative engine works differently. It splits the web into passages, indexes them, then, faced with a request, it manufactures several sub-queries and fetches the passages closest to each one. That fan-out mechanism is described in detail in our article on query fan-out. At the moment of retrieval, what gets compared is a passage, not a position in a hierarchy.
Two consequences follow from that mechanism alone. First: a deep page is not penalised for its depth when a passage is retrieved, provided it was indexed in the first place. Second: architecture becomes decisive upstream, because a page no robot reaches in the served HTML never enters the index that gets queried.
Which is where the measured result above regains all its weight. The subject is not the shape of the tree, it is whether pages exist at all for a robot that renders nothing. How to declare your perimeter to AI engines is covered in our llms.txt guide, and the full approach on our AI Search architecture page.
Our own numbers, the bad ones included
vydera.com sits in the panel, measured exactly like the others, and it does not come out ahead everywhere.
- 52.7% of pages reachable from the homepage, against a 58.6% median. Below the panel median.
- 8.47 content links per page, against a median of 8.74. Just under, again.
- 74 glossary pages out of reach, 37 French and 37 English, because the page listing them receives no content link. That is our own language switcher problem, only worse: the hub lives in the nav.
- Two figures in our favour: 4.8% orphan pages and 40.8% concentration, the second lowest in the panel behind agence-slashr.fr.
A control run covered the /fr perimeter alone: 52.4% reachable against 52.7% site-wide, 4.8% orphan in both readings. The localised measurement tells the same architectural story as the global one.
A third reading, run the same day with an earlier script, counted 825 links on the /fr perimeter where the new one counts 699. The difference comes down to one rule: the new script strips nav blocks nested inside the main tag, the old one did not. Same site, same day, a 126-link gap. That is the shortest possible demonstration of the caveat stated above.
Placing your own site, and running the measurement again
The benchmarks below come from the nine comparable sites. They place an order of magnitude, they do not grade a site.
The measurement replays with no account and no paid tool, on public HTML: read the sitemap, open every URL, extract the links in the content zone, walk the graph breadth first from the homepage. Three precautions are worth repeating verbatim.
- Do not half-crawl. A page looks orphan the moment you have not opened the page that links it. Past a ceiling, a site should leave the comparison rather than be measured partially. Three panel sites brush ours, set at 1,600 URLs: keyweo.com at 1,589, digidop.fr at 1,438, uplix.fr at 1,285.
- Watch for 429 codes. natural-net.fr rate-limited on the first pass, at three parallel requests. The whole measurement was rerun at one request every 1.5 seconds, and that second reading is what produces the 691 pages, 46.5% orphans and 7.26 links per page published here. The first pass is not published: it described a throttle, not a site. A cluster of 429s is not an architecture, it is a site you did not measure.
- Publish your content-zone definition. Without it your number compares to nothing, not even to your own from last month.
One last useful marker: the share of links pointing at a page absent from the sitemap varies wildly, from 41 on vydera.com to 930 on natural-net.fr and 1,507 on digidop.fr. Those links are excluded from the graph, which slightly overstates dead ends on sites with incomplete sitemaps.
So, does siloing hold up?
Nothing here condemns semantic siloing, and nothing rescues it. What nine sites measured on the same day show is that its signature is not readable in the HTML of the sector that sells it, and that the architecture problem staring back is far more mundane: entire chunks of site the homepage never reaches, because the link lives in a menu.
Four things to do before redrawing anything:
- Measure the share of pages the homepage reaches through content links, served HTML, no JavaScript.
- Separate orphans from unreachables. Two different failures, two different fixes.
- Look at what the unreachable bucket actually holds. In this panel it is almost always a language, a paginated index or a hub that only lives in the nav.
- Push one content link toward each of those buckets, from a page that is already reachable. That is half a day of work, not a rebuild.
For the definitions: internal linking and crawl.
What exactly is semantic siloing?
A content tree where each page covers one narrow sub-topic and links only to its immediate neighbours in the hierarchy, so link equity and relevance concentrate on a target page. Two measurable properties follow: shallow depth from the homepage, and heavy concentration of inbound links on a few head pages.
Has semantic siloing become pointless with AI engines?
Our measurement cannot settle that: no AI engine retrieval was measured. What it does show is that across the 9 French SEO sector sites measured on 27 August 2026, shallow depth paired with heavy concentration never appears (Spearman's rho of -0.289 across 9 sites, +0.025 across 8). The shape of the tree matters less than whether a page is reachable in the served HTML.
What is the difference between an orphan page and an unreachable page?
An orphan page receives no content link. An unreachable page receives some, but no link path connects it to the homepage. The gap is spectacular on seo.fr: 0.4% orphan pages and 41.4% unreachable ones, including 204 English pages that link to each other without ever connecting back to the homepage.
How many internal links per page do SEO agency sites carry?
Across the 9 sites measured on 27 August 2026, content links per page run from 3.02 to 18.74, with a median of 8.74. Nav, header and footer links are excluded. One warning: the figure depends entirely on how the content zone is detected, to the point of varying by a factor of 6 on a single site depending on the rule applied.
How many clicks deep should a page sit from the homepage?
Exactly one panel site meets the three-click rule strictly, and it is the degenerate case whose homepage reaches only 6 pages. The other eight run from 4 to 17 clicks at maximum depth, and 2 to 7 at median depth. The useful question is not click count, it is the share of pages the homepage reaches: a median of 58.6%, meaning 41.4% of the site sits outside the graph for a robot that runs no JavaScript.
How do I measure my own internal link architecture?
Read the sitemap, open every URL without sampling, extract the links in the content zone outside nav, header and footer, then walk the graph breadth first from the homepage. Two traps: a partial crawl flags pages as orphan when they were simply never opened, and a cluster of 429 codes manufactures a false profile. One site in our panel rate-limited on the first pass: the measurement was rerun end to end at one request every 1.5 seconds, and only that second reading is published.




