To decode Cloudflare email protection, take the hex string from the data-cfemail attribute or from the part after # in a /cdn-cgi/l/email-protection link. The first byte is the key. XOR every following byte with it and read the result as UTF-8: that is the address. On 7 October 2026 this one rule decoded 504 of 508 payloads on 54 real pages.
The rest of this guide shows the format, working code in Python and JavaScript, what we found when we ran it on real contact pages, and when it is simpler to let an API do it. If you only have one address to check, paste it into the free Cloudflare email decoder.
What Cloudflare does to an email address
Email Address Obfuscation is a Cloudflare setting that lives under Scrape Shield. When a site turns it on, Cloudflare's edge rewrites the HTML on its way out. Every address written as text is replaced by a placeholder, every mailto: link is pointed at a Cloudflare path, and a small script called email-decode.min.js is added so that a browser with JavaScript can put the addresses back.
<!-- what the site wrote -->
<a href="mailto:info@example.com">info@example.com</a>
<!-- what Cloudflare serves -->
<a href="/cdn-cgi/l/email-protection#5a33343c351a3f223b372a363f74393537">
<span class="__cf_email__" data-cfemail="5a33343c351a3f223b372a363f74393537">[email protected]</span>
</a>
A visitor in a browser never notices. Anything that reads the raw HTML does: curl, Python requests, most crawlers and the agents that fetch pages for AI assistants all get the literal placeholder, and the link leads to a /cdn-cgi/ path with no page behind it. On 7 October 2026 even Google's result snippets for Italian dental practices printed the placeholder where the address should have been.
The format: one key byte, then the address XOR the key
The payload is hexadecimal, two digits per byte. The first byte is a key between 0 and 255 that Cloudflare picks for each occurrence. Every byte after it is one byte of the address, XOR the key. Take the payload above, 5a33343c351a3f223b372a363f74393537:
| Byte | XOR 0x5a | Character |
|---|---|---|
| 5a | the key (90) | none |
| 33 | 0x69 | i |
| 34 | 0x6e | n |
| 3c | 0x66 | f |
| 35 | 0x6f | o |
| 1a | 0x40 | @ |
| 3f 22 3b and the rest | 0x65 0x78 0x61 and the rest | e x a and the rest |
Sixteen bytes after the key spell the sixteen characters of the example address. The decode script reads the bytes as UTF-8, so an address with non-ASCII characters comes out right as long as you do the same.
Two details catch people out. First, the key changes from one occurrence to the next, so the same address appears under different hex strings on the same page: decode first, then remove duplicates. Second, some sites hide the address before Cloudflare ever sees it, with HTML entities (in…), percent-encoding (%40 for the at sign, %20 for a stray space) or a mailto: prefix with a ?subject= tail. Cloudflare encodes whatever is there, so after the XOR you may still have a layer to remove.
Decode it in Python
Standard library only. The function returns the bare address, or None when the payload holds something else, such as a share link with a subject and no recipient.
import html
import re
from urllib.parse import unquote
EMAIL = re.compile(r"[\w.+-]+@[\w-]+(?:\.[\w-]+)+")
def decode_cfemail(hex_str):
data = bytes.fromhex(hex_str)
key = data[0]
text = bytes(b ^ key for b in data[1:]).decode("utf-8", "replace")
text = unquote(html.unescape(text)) # layers the site added first
m = EMAIL.search(text)
return m.group(0) if m else None
def emails_in_html(page):
payloads = re.findall(r'data-cfemail="([0-9a-fA-F]+)"', page)
payloads += re.findall(r"/cdn-cgi/l/email-protection#([0-9a-fA-F]+)", page)
found = {decode_cfemail(p) for p in payloads}
return sorted(e for e in found if e)
print(decode_cfemail("5a33343c351a3f223b372a363f74393537")) # info@example.com
If you parse with BeautifulSoup, the selector is [data-cfemail] for the attribute and a[href*="/cdn-cgi/l/email-protection"] for the links. Replace the placeholder text with the decoded address before you extract text, so that your Markdown or CSV never contains [email protected].
Decode it in JavaScript
The same rule in a browser or in Node. TextDecoder handles the UTF-8 step.
function decodeCfemail(hex) {
const bytes = hex.match(/../g).map((h) => parseInt(h, 16));
const key = bytes[0];
const raw = Uint8Array.from(bytes.slice(1), (b) => b ^ key);
let text = new TextDecoder().decode(raw);
text = text.replace(/&#(\d+);/g, (_, d) => String.fromCharCode(+d));
try { text = decodeURIComponent(text); } catch {}
const m = text.match(/[\w.+-]+@[\w-]+(\.[\w-]+)+/);
return m ? m[0] : null;
}
document.querySelectorAll("[data-cfemail]").forEach((el) => {
el.textContent = decodeCfemail(el.dataset.cfemail);
});
This is close to what Cloudflare's own script does in the page, minus the extra layers. the decoder tool runs the same code with a step-by-step view of the bytes, if you want to check a payload by eye.
What we found on 54 real pages
A format is easy to describe. What matters for a parser is what it meets in the wild, so on 7 October 2026 we measured it. We ran four Google searches for pages whose snippet showed [email protected]: Italian dental practices, law firms, hotels and US dentists. Of the 68 result pages, 54 returned HTML to a plain HTTP request, and we decoded every payload in that HTML with the Python rule above.
| Measure, 7 October 2026 | Result |
|---|---|
| Pages still serving Cloudflare email protection | 49 of 54 |
| Encoded payloads on those pages | 508 |
| Payloads in a data-cfemail attribute | 269 |
| Payloads in a link with #hex | 239 |
| Payloads that decoded to an address | 504 of 508 |
| Share links with no recipient (subject or body only) | 4 |
| Addresses percent-encoded by the site before Cloudflare | 3 |
| Pages with the same address under two or more keys | 40 of 49 |
| Distinct addresses recovered, summed per page | 159 |
Three things follow. Every one of the 49 protected pages used the data-cfemail attribute, and 39 of them also used #hex links, so a parser needs both. Most pages repeat an address under different keys, which is why you deduplicate after decoding. And a few payloads are not addresses at all: share buttons built as mailto:?subject=…&body=… get encoded too, and a decoder should return nothing for them rather than garbage.
The same morning we sent the contact page of an Italian lighting manufacturer, which protects every address on the page, through one call to our scraping API. It came back on the plain HTTP tier in 7.3 seconds with 59 addresses in metadata.contactEmails, each one marked as decoded from Cloudflare.
When a scraper should do it for you
Decoding one payload is a curiosity. Decoding them across thousands of pages means also fetching those pages, which is the harder half: residential exits, retries, and a browser only when a page truly needs one. The web scraping API does the whole job in one call and decodes Cloudflare-protected addresses natively, with no browser and no JavaScript. It covers:
data-cfemailon aspan, anaor any other tag;/cdn-cgi/l/email-protection#…links, relative or absolute, and links without the#hex(the decoded label is used);- the protection URL written as plain text;
- the layers sites add first: HTML entities,
%40, the mailto prefix and the subject tail.
Markdown and text output carry the address with a mailto link. On zkdental.it the link went from the first line below to the second:
[info@zkdental.it](https://www.zkdental.it/cdn-cgi/l/email-protection#7355…)
[info@zkdental.it](mailto:info@zkdental.it)
With format: "html" the page comes back byte for byte as served, except the placeholders. Every response also lists the page's own contacts: metadata.contactEmails with a source of json-ld, cloudflare or mailto (in that order, without duplicates), and metadata.contactPhones from JSON-LD and tel: links.
curl https://api.quanticdata.io/v1/scrape \
-H "Authorization: Bearer $QD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://www.reggiani.net/it/contatti/", "format": "markdown"}'
# payload.metadata (abridged)
"contactEmails": [ { "email": "contact@reggiani.net", "source": "cloudflare" }, … ],
"contactPhones": [ { "phone": "00390297070340", "source": "tel" } ]
A page costs $0.0002 on the plain HTTP tier and $0.001 when a browser render is needed, failed pages are not billed, and every key comes with $2 of free API usage per month. If you start from a list of domains rather than URLs, the email scraper API returns one contact record per site. The full request reference is in the API documentation, and the plans are on the pricing page.
Limits, law and the B2B line
The obfuscation is not encryption. The key travels with the payload, the algorithm is public, and the decode script ships with every protected page, so decoding it breaks no lock. It does not change what you may do with the result.
- Personal data stays personal data. In the EU, GDPR applies to a named person's work address as much as to a private one. You need a lawful basis, usually legitimate interest documented for B2B contact, an opt-out in every message and a way to honour erasure requests.
- Use published business contacts for business purposes. The common legitimate case is the front-door address a company prints on its contact page, used to propose something relevant to that company. Consumer data and bulk mailing without a basis are out.
- Respect the site. Read the terms, keep request rates modest and do not use addresses for anything the page clearly did not publish them for. For the wider picture of where companies publish their addresses, see our guide.
On the site-owner side the advice is simpler: if the 404 links to /cdn-cgi/l/email-protection in your SEO crawler bother you, switch the feature off in the Cloudflare dashboard or exclude single blocks with the email_off comment tags. The protection stops some low-effort harvesters; it does not stop anyone who has read this far.