Free tool · runs in your browser
Email extractor
Paste any text or add a CSV or TXT file. The email extractor finds every email address, removes duplicates and sorts the list, ready to copy, download as CSV or validate.
Extracted addresses
Paste text on the left or add a file. The addresses appear here as you type.
How to extract email addresses from text
- Paste or add files. Paste an email thread, a web page’s text or source, a signature block or a CRM export into the box, or add CSV, TXT, vCard, HTML or JSON files (up to 20 MB together). You can combine both.
- Read the counts. The tool shows how many addresses it found, how many are unique and how many duplicates it removed. Matches that break the format rules are listed separately.
- Choose the options. Sort A to Z, group by domain or keep the order of the text; keep only some domains or exclude them; keep duplicates if you need every occurrence.
- Copy or export. Copy the list, or download it as CSV or TXT.
- Validate the list. A found address is only text shaped like an address. Check the mailboxes before you send (see below).
The list updates as you type, so you can also paste a list you already have just to de-duplicate and sort it.
What the email address extractor finds
The email address extractor looks for addresses wherever they sit: in headers, quotes, angle brackets, CSV cells, HTML attributes and mailto: links. It lowercases every address and drops punctuation that belongs to the sentence around it.
| In the text | Result |
|---|---|
"Jane Doe" <Jane | jane.doe@example.com |
mailto: | press@example.com |
Write to sales@example | sales@example.co.uk (the final period is dropped) |
o'brien@example.ie | o'brien@example.ie |
jane+news@example.com | jane+news@example.com (plus tags are kept) |
billing [at] example [dot] org | billing@example.org (with decoding on) |
jane@example.com or mailto: | jane@example.com (with decoding on) |
jane..doe@example.com | skipped: two dots in a row |
logo@2x.png | skipped: a file name, not an address |
name@localhost, jane@example | not found: no top-level domain |
josé@example.com, jane@bücher.example | not found: non-ASCII characters |
Internationalized domains are found in their ASCII form (jane@xn--bcher-kva.example). Addresses with non-ASCII characters in the local part need SMTPUTF8 support on every server on the way, so they are left out, as in the email regex guide. The rules behind each case are in the guide to the email address format.
Decoding only covers forms that can’t be ordinary words: [at], (at), {at} and <at> (the same for dot), the HTML entities for @ and URL-encoded %40. Plain jane at example dot com is not decoded, because “at” and “dot” also appear in normal sentences.
Remove duplicate emails and filter by domain
Duplicates. Addresses are compared after lowercasing, so Jane@Example.com and jane@example.com count as one. Strictly, RFC 5321 says the part before the @ may be case-sensitive, but it also says exploiting that “impedes interoperability and is discouraged”, and domain names are never case-sensitive. Turn on Keep duplicates to list every occurrence instead.
Domain filter. Choose “Only these domains” or “Exclude these domains” and enter one or more domains, separated by commas or spaces. A domain also matches its subdomains: example.com covers mail.example.com, but not notexample.com. Typical uses: keep only the addresses at a customer’s domain, or drop your own colleagues from a thread.
Group by domain. Sorts the list by domain, then by address, and shows how many addresses each domain has.
CSV export. The file has three columns: email, domain and occurrences (how often the address appeared in your input). Values that start with =, +, - or @ get a leading apostrophe, so spreadsheet apps don’t treat them as formulas. The TXT download has one address per line.
Validate the extracted list
Extraction tells you that something looks like an address, not that it receives mail. Old threads and exports are full of addresses that bounce now: people leave companies, domains expire, and typos survive copy and paste. Sending to them raises your email bounce rate.
Click Copy list and open the bulk verifier under the results, then paste the list into the bulk email verifier. The free tool takes the first 50 addresses of a list and, without an account, checks a few addresses an hour. For a whole export, use the email list cleaning service: it takes a CSV or a JSON array of up to 100,000 addresses, checks every unique address in the background, and doesn’t charge for duplicates or rows that aren’t valid addresses. The free plan includes 100 validations a month.
The regex behind the extractor
The extractor uses the search pattern from our email regex guide, extended for apostrophes, punycode top-level domains and length limits:
/(^|[^A-Za-z0-9._%+'@-])([A-Za-z0-9._%+'-]{1,64}@[A-Za-z0-9.-]{1,252}\.(?:xn--[A-Za-z0-9-]{2,59}|[A-Za-z]{2,63}))(?![A-Za-z0-9@])/g
| Part | What it does |
|---|---|
(^|[^A-Za-z0-9._%+'@-]) | The address starts at the beginning of the text or after a character that can’t be part of it, so jane@doe@example.com isn’t cut into a fake address. |
[A-Za-z0-9._%+'-]{1,64} | The local part: letters, digits, dots, %, +, ', _ and -, at most 64 characters. |
[A-Za-z0-9.-]{1,252}\. | The domain labels. |
(?:xn--…|[A-Za-z]{2,63}) | A top-level domain of letters, or an internationalized one in punycode. |
(?![A-Za-z0-9@]) | The address must end here, so jane@example.com123 isn’t shortened to a different address. |
The upper limits are there for speed as much as for correctness: without them, a long run of letters in a large file can make a regex engine try millions of combinations. Each match is then checked against the practical validation pattern from the same guide. Matches that fail it, such as jane.@example.com or jane@-example.com, are shown under “Skipped matches” instead of being dropped silently. Leading dots and apostrophes that belong to the surrounding text ('jane@example.com', ...jane@example.com) are removed first.
To check single addresses against the same rules, use the email syntax checker.
Files, limits and privacy
- Everything stays in your browser. Pasted text and files are read with the browser’s FileReader and processed on your device. Nothing is uploaded, and the page doesn’t save the text or the list.
- No crawling. The tool never fetches or scans other websites. It works on text you give it.
- File types. Plain text formats: CSV, TSV, TXT, vCard, HTML, XML, JSON, Markdown, EML and LDIF. Save Excel, Numbers, Word or PDF files as CSV or TXT first. UTF-8 is read by default; files that start with a byte order mark (such as UTF-16 exports) are decoded accordingly.
- Size. Up to 20 MB of files at once. Very large pastes work too, but files are faster, since the browser doesn’t have to show the text.
Use extracted addresses responsibly
Extracting addresses from your own mail, exports and documents is everyday list work. Mailing them is where the rules start:
- EU: consent for marketing email. The ePrivacy Directive allows email for direct marketing only to recipients who “have given their prior consent” (Directive 2002/58/EC, Article 13). The exception covers a company’s own customers, whose details it got in the context of a sale, and only for its own similar products or services: they must be able to object easily and free of charge, both when the details are collected and in every message.
- EU: the GDPR applies to the list itself. An address that identifies a person is personal data (GDPR, Article 4(1)). Processing it needs a legal basis (Article 6(1)), and when you collect personal data from somewhere other than the person, you must inform them, at the latest when you first contact them (Article 14(3)(b)).
- US: CAN-SPAM. The law doesn’t require consent, but every commercial email needs a working opt-out, your physical postal address and honest headers and subject lines, and opt-outs must be honored within 10 business days. The FTC’s compliance guide lists penalties of up to $53,088 per email. If a message breaks these rules and went to an address collected by automated means from a website that says it doesn’t share its addresses, 15 U.S.C. § 7704(b)(1) makes it an aggravated violation.
- Deliverability. Spamhaus seeds spam trap addresses in places online; mail to them shows that a sender is scraping addresses from the web or buying lists from someone who does (Spamhaus, February 2022). The guide to spam traps explains what happens next.
This is not legal advice. If you need contact details for a specific person, the email finder and the guide on how to find someone’s email address are better starting points than a scraped list.
Sources
Checked October 9, 2026: RFC 5321, section 2.4; Directive 2002/58/EC; Regulation (EU) 2016/679 (GDPR); FTC: CAN-SPAM Act compliance guide; 15 U.S.C. § 7704; Spamhaus: Spamtraps – fix the problem, not the symptom; IANA root zone TLD list (for the file-name rule: .png and .jpg are not top-level domains, .zip and .mov are).
Frequently asked questions
How do I extract emails from a text or document?
How do I remove duplicate emails from a list?
Can the email extractor scan a website?
Is extracting emails legal?
Bulk email verification API
Validate the whole list, not just the first 50
Upload the CSV or send a JSON array. The API checks every unique address in the background, can call a webhook when the job has finished, and returns your list with a result for every row.
/v1/bulk Email list cleaning Read the documentation
One request per unique, valid address. 100 free requests every month, no credit card required.
POST https://api.emailvalidation.io/v1/bulk
{
"job_id": "629954218123464704",
"status": "queued",
"source": "csv",
"counts": {
"total": 5,
"unique": 3,
"duplicates": 1,
"invalid": 1,
...
},
"quota": {
"charged": 3,
"refunded": 0
},
...
}