TL;DR
- One export takes at most 50,000 rows, so 1M contacts is 20 exports. Split the audience into slices with filters, then loop.
- It runs for about two days, because each account may export 500,000 rows a day. Your own time is about 45 minutes (both are estimates).
- It spends about 1,000,000 LeadOcean records, one per row written. Pro is $499 a month flat. Free has 1,000 records, enough for a dry run.
This guide builds the loop in Python. The bulk export is also on the Exports page of the app if you would rather click than code. New to the idea? Read what a bulk export is first.
Prerequisites
- A LeadOcean account on Pro. Free stops at 1,000 records, so a 1M pull needs Pro. Free still runs every step below with a small
limit. - An API key with the
searchandenrichscopes, exported as$LEADOCEAN_API_KEY. Thefullcolumn preset reveals emails and phones, so it needsenrich. - Python 3.9 or later and
pip install requests. - Disk space for the files. A 1M-row CSV with contact columns is large.
Read the export docs once before you start (checked September 2026).
Step 1: Check your key and the ceilings
Start with two free calls. The first lists the export ceilings. The second shows where your account stands.
curl -s https://api.leadocean.io/v1/exports/columns | python3 -c \
"import json,sys; print(json.load(sys.stdin)['data']['ceilings'])"
curl -s https://api.leadocean.io/v1/usage -H "x-api-key: $LEADOCEAN_API_KEY"The ceilings are 50,000 rows per export, 500,000 a day and 10,000,000 per billing period. One export runs at a time per account. The rest queue.
Step 2: Size the audience for free
Send count=true as a query parameter with limit: 1 in the body. It returns the total in meta.total and costs no records.
curl -s -X POST "https://api.leadocean.io/v1/people/search?count=true" \
-H "x-api-key: $LEADOCEAN_API_KEY" -H "content-type: application/json" \
-d '{
"country": ["US"],
"jobLevel": ["Director"],
"emailStatus": ["verified", "catch_all_valid", "catch_all"],
"limit": 1
}'Pass emailStatus yourself so the count and the export use the same definition of mailable. meta.total is capped at 100,000 on the API, so a larger audience reads as a floor.
Step 3: Split it into slices under 50,000 rows
A single export always starts at the top of its filter set. Ten exports with the same filters give you the same 50,000 people ten times. You need slices that do not overlap.
Filters that take one value per person make clean slices: jobFunction (22 values) crossed with employeeRange (8 brackets). Count each cell for free and keep the cells that fit. If a cell is over 50,000, split it again by emailStatus.
"""plan.py: build non-overlapping slices worth 1,000,000 rows"""
import json, os, requests
BASE = "https://api.leadocean.io"
H = {"x-api-key": os.environ["LEADOCEAN_API_KEY"], "content-type": "application/json"}
TARGET, CEILING = 1_000_000, 50_000
STATUSES = ["verified", "catch_all_valid", "catch_all"]
FIXED = {"country": ["US"], "jobLevel": ["Director"]}
FUNCTIONS = ["Advertising & Marketing", "Engineering", "Finance & Accounting",
"General Business & Management", "Human Resources", "Information Technology",
"Operations", "Sales & Business Development"] # add the other 14 as needed
RANGES = ["1-10", "11-50", "51-200", "201-500", "501-1000", "1001-5000", "5001-10000", "10001+"]
def count(filters):
body = {**filters, "limit": 1}
r = requests.post(f"{BASE}/v1/people/search", headers=H, params={"count": "true"}, json=body)
r.raise_for_status()
meta = r.json()["meta"]
return meta.get("total") or 0
slices = []
for fn in FUNCTIONS:
for rng in RANGES:
cell = {**FIXED, "jobFunction": [fn], "employeeRange": [rng]}
cell_total = count({**cell, "emailStatus": STATUSES})
if cell_total == 0:
continue
parts = [{**cell, "emailStatus": STATUSES}] if cell_total <= CEILING else \
[{**cell, "emailStatus": [s]} for s in STATUSES]
for f in parts:
n = count(f)
if n == 0:
continue
if n > CEILING:
print("still over 50,000, split further:", f)
continue
slices.append({"filters": f, "rows": n})
total = 0
for s in slices:
s["rows"] = min(s["rows"], TARGET - total)
total += s["rows"]
if total >= TARGET:
slices = slices[: slices.index(s) + 1]
break
json.dump(slices, open("slices.json", "w"), indent=1)
print(len(slices), "slices,", total, "rows")If the loop prints "still over 50,000", add one more filter to that cell, such as seniority or continent, and rerun. Every slice must count at or under 50,000.
Step 4: Start each export and poll until it is done
Start one export at a time. The daily ceiling counts rows already reserved, so twenty parallel starts would fail on the eleventh. When the day is full the API answers 429 with details.resetsAt. The script sleeps until then.
"""run.py: one export per slice, download each file"""
import json, os, time, requests
from datetime import datetime, timezone
BASE = "https://api.leadocean.io"
H = {"x-api-key": os.environ["LEADOCEAN_API_KEY"], "content-type": "application/json"}
def start(i, s):
body = {"name": f"slice {i:03d}", "filters": s["filters"],
"limit": s["rows"], "preset": "full"}
while True:
r = requests.post(f"{BASE}/v1/exports", headers=H, json=body)
if r.status_code == 202:
return r.json()["data"]
err = r.json().get("error", {})
if r.status_code == 429:
when = (err.get("details") or {}).get("resetsAt")
wait = 60
if when:
t = datetime.fromisoformat(when.replace("Z", "+00:00"))
wait = max(60, (t - datetime.now(timezone.utc)).total_seconds())
print(f"slice {i}: {err.get('message')} Sleeping {int(wait)}s")
time.sleep(wait)
continue
raise SystemExit(f"slice {i}: {r.status_code} {err}")
def wait_done(export_id):
while True:
d = requests.get(f"{BASE}/v1/exports/{export_id}", headers=H).json()["data"]
if d["status"] in ("done", "failed", "cancelled"):
return d
time.sleep(d.get("pollAfter") or 10)
for i, s in enumerate(json.load(open("slices.json"))):
path = f"slice_{i:03d}.csv"
if os.path.exists(path):
continue # already downloaded on an earlier run
job = start(i, s)
print(f"slice {i}: {job['id']} perRow={job['perRow']} remaining={job['remaining']}")
d = wait_done(job["id"])
if d["status"] != "done":
print(f"slice {i}: {d['status']} {d.get('error')} rows={d['rows']} billed={d['billed']}")
continue
with requests.get(f"{BASE}/v1/exports/{job['id']}/download", headers=H, stream=True) as r:
r.raise_for_status()
with open(path, "wb") as f:
for chunk in r.iter_content(1 << 20):
f.write(chunk)
print(f"slice {i}: {d['rows']} rows, billed {d['billed']}")Check perRow on the first 202 and multiply it by 1,000,000 before you run the rest. The docs say it is 1. Files are kept for 7 days, so download each one the day it finishes.
Step 5: Merge and dedupe
The slices do not overlap, but rerunning a slice can. Dedupe on person_id while you merge, and stream the rows so memory stays flat.
"""merge.py: merge the slice files and dedupe on person_id"""
import csv, glob
seen, writer = set(), None
with open("contacts_1m.csv", "w", newline="") as out:
for path in sorted(glob.glob("slice_*.csv")):
with open(path, newline="") as f:
reader = csv.DictReader(f)
if writer is None:
writer = csv.DictWriter(out, fieldnames=reader.fieldnames)
writer.writeheader()
for row in reader:
if row["person_id"] in seen:
continue
seen.add(row["person_id"])
writer.writerow(row)
print(len(seen), "unique people")Compare that number with the sum of rows from the export jobs. They should match.
What you get
These are free counts from the LeadOcean MCP tool leadocean_count_leads, run on 2026-09-30. Each uses the mailable default: verified, catch_all_valid or catch_all.
| Filters (country US, jobLevel Director) | Mailable people |
|---|---|
| all job functions | 4,494,355 |
| jobFunction Sales & Business Development | 294,762 |
| + employeeRange 1-10 | 23,392 |
| + employeeRange 11-50 | 35,202 |
| + employeeRange 51-200 | 32,835 |
| + employeeRange 201-500 | 25,098 |
| + employeeRange 10001+ | 84,673 |
The 10001+ cell is over 50,000, so Step 3 splits it by emailStatus. The other cells fit in one export each.
A started export returns this shape (response shape from the API reference, placeholder values):
{
"success": true,
"data": {
"id": "<export id>",
"name": "slice 000",
"status": "queued",
"requested": 50000,
"estimatedTotal": 50000,
"reveal": "email_phone",
"perRow": 1,
"remaining": null,
"pollAfter": 10,
"file": null
}
}The full preset writes 45 columns. They cover identity, profile, address, contact flags, up to three emails, two phones, current role, education and skills. The search preset writes 27 and needs no enrich scope.
Troubleshooting
| Error | Cause | Fix |
|---|---|---|
| 401 | Key missing, wrong or revoked | Send the key in x-api-key. Create a new one in the app if it was revoked. |
| 429 above 100 requests a second | You passed the per-key rate | Wait for Retry-After. The plan loop is sequential; add time.sleep(0.05) if it trips the limit. |
429 export_limit_daily | The 500,000-row day is full, counting rows still reserved | Sleep until details.resetsAt, which run.py does. |
| 402 | limit is more than the records you have left | Free: upgrade, the 1,000 do not reset. Pro: write to support@leadocean.io before a pull this size. |
400 export_limit_rows | limit is above 50,000 | Cap each slice at 50,000 rows. |
| 0 rows or 0 matches | A filter value is not in the catalogue, such as US typed as USA | Look the value up with GET /v1/enums/{name} or leadocean_list_enum_values, then rerun the count. |
A 403 means the key lacks the enrich scope for the full preset. Use a key with both scopes or switch to search.
FAQ
How many exports does 1 million rows take?
At least 20, because each export stops at 50,000 rows. In practice you run more, because slices rarely land at exactly 50,000. Step 3 prints the real number.
Can I run all the exports at once?
No. One export runs per account at a time, and the daily ceiling counts reserved rows. Start them in sequence, as run.py does.
What does 1M rows cost on Pro?
Pro is $499 a month flat, with no per-record price and no overage invoice (pricing). Past the fair-usage line a very large month is paced down to one request a minute until the reset. Email support before a 1M pull to agree the headroom.
Why not page through search instead?
One search walk stops at 10,000 rows, and search pages hold at most 100 records. Bulk export is the route for big pulls. The hub page covers when to use which.
Can I raise the 50,000 and 500,000 ceilings?
Yes, for an account. Write to support@leadocean.io (export docs, September 2026).
Test the 1M pipeline on 1,000 free records
Free to start. No credit card. 1,000 records to spend whenever you like.
Get your free API key →