An XML sitemap is not just a list of URLs generated automatically by a CMS. On a Parisian site whose content evolves with local events, updates to services, or neighborhood news, the quality of the file directly affects how quickly search engines discover new pages and ignore old ones.
Reliability of lastmod: the only signal that matters in a sitemap
We often observe sitemaps where the lastmod tag displays the file generation date for each URL, with no connection to the editorial reality. Google has clarified its position: it uses this date only when it is consistently and verifiably accurate, meaning it reflects a significant change in the main content, structured data, or links.
Changing the copyright year in the footer does not constitute a significant update. On a Parisian site where hours, prices, or event pages change regularly, lastmod should come from the actual editorial history, not from the date the file was rebuilt.
A CMS like WordPress, paired with a properly configured SEO plugin, can timestamp each URL based on the last effective save of the article. In contrast, a custom script that regenerates the sitemap every night by applying the current date to all entries sends a mixed signal. Robots eventually ignore the dates altogether, negating the file’s usefulness for targeted recrawling.
By analyzing the sitemap of the Paris Astuce site, we see a structure that reflects this principle: the URLs are organized by content type, with dates consistent with actual publications.

Discrepancies between sitemap, CMS, and Search Console on a Parisian website
The comparison between the URLs in the sitemap, the pages actually published by the CMS, and those indexed in Google Search Console constitutes an audit that most consumer guides overlook. These discrepancies, however, reveal concrete problems.
Pages present in the sitemap but absent from the index
A sitemap is a discovery signal, not an imperative crawl request. A URL listed in the file may remain excluded if it has a noindex directive, if its content is deemed too weak, or if no internal link points to it. On a Parisian site with dozens of local service pages, we recommend regularly cross-referencing the “Pages” report in Search Console with the sitemap content to identify these exclusions.
Indexed pages but absent from the sitemap
The opposite case is equally common. Test pages, drafts mistakenly made public, or URLs with tracking parameters can be crawled via internal linking without ever appearing in the sitemap. A clean sitemap contains only the canonical URLs to be indexed. Any URL you do not want to see in search results has no place in this file.
Here are the checks to perform periodically:
- Extract the complete list of URLs from the sitemap and compare it with the export of published pages from the CMS (status “published” only, excluding drafts and private pages)
- Check in Search Console that each URL in the sitemap is either indexed or deliberately excluded by an explicit directive (noindex, canonical to another page)
- Identify orphaned URLs indexed by Google but absent from the sitemap, which often signal outdated pages or uncleaned technical routes
Organization of sitemaps for a multi-section Parisian website
A Parisian website covering multiple themes (districts, outings, good deals, news) benefits from segmenting its sitemap rather than grouping everything into a single file. The sitemap index mechanism allows for referencing multiple child files, each dedicated to a type of content.
This segmentation is not just aesthetic. It allows for observing the indexing rate by category in Search Console. If the “events” pages show a significantly lower indexing rate than the “neighborhood guides” pages, the issue likely lies in the freshness or depth of the event content, not in the technical setup.

Naming and structuring child files
We recommend explicit naming: sitemap-posts.xml, sitemap-pages.xml, sitemap-categories.xml. This convention facilitates reading Search Console reports and speeds up diagnosis in case of an indexing drop in a specific segment.
Each child file should contain only URLs returning an HTTP 200 code. 301 redirects, 404 errors, and soft 404 pages pollute the file and waste crawl budget. On a site whose content is related to Parisian news, expired pages (past events, closed venues) should be removed from the sitemap as soon as they are no longer relevant.
Interaction between robots.txt and XML sitemap
The robots.txt file can declare the location of the sitemap via the Sitemap directive, which remains the simplest method for robots to find it without manual submission. However, an inconsistency between the two files creates a silent conflict.
If robots.txt blocks an entire directory via Disallow and the sitemap contains URLs from that directory, search engines receive a contradictory signal. The sitemap never overrides a Disallow rule in robots.txt. The URL will be discovered but not crawled, generating errors in Search Console without any SEO benefit.
On a Parisian WordPress site, this problem frequently occurs with image directories, author pages, or tag archives that some SEO plugins include by default in the sitemap while blocking them in robots.txt. A cross-audit of the two files avoids these contradictions.
The sitemap of a Parisian website is not a file to be generated once and then forgotten. It is a continuous diagnostic tool whose value entirely depends on the rigor with which it reflects the actual state of the site, URL by URL, date by date.



