How to scrape Realtor.com listings without the 429 soft-block
Realtor.com has no self-serve listing API, but each search results page carries its listings as JSON in a __NEXT_DATA__ script tag. Plain HTTP clients get blocked quickly, and heavy traffic from one address triggers a 429 page that reads "This is taking longer than usual". Load search pages in a real browser, read that JSON, and give each browser its own static US residential ISP proxy at a slow, steady pace. Stat Proxies static ISP IPs cost $2.50 per IP per month with unlimited bandwidth and a 25 IP minimum.
At a glance
- Recommended
- Static ISP Proxies
- Blocker
- 429 soft-block page and fast blocks on plain HTTP clients
- Client
- Real browser (Playwright) so the page hydrates
- Data location
- __NEXT_DATA__ JSON on search and listing pages
- Search URLs
- /realestateandhomes-search/<City>_<ST>/pg-<n>
- Bandwidth
- Unmetered on every plan
- Starting at
- $2.50 / proxy / mo
Why Realtor.com scrapers break.
Realtor.com's data is easy to read once a page has loaded. Most failures come from the client and the pace, not the parsing.
- Plain HTTP libraries like requests or urllib get blocked fast. The site looks at more than the IP, so a request that does not look like a browser stands out on its own.
- Push too many pages through one address and you get a 429 page with the message "This is taking longer than usual". It is a soft rate limit, so retrying straight away from the same IP keeps it going.
- The raw HTML from a non-browser fetch may not include the hydrated listing JSON. A scraper that only checks for status 200 can store empty results without noticing.
- Cloud and data centre IP ranges carry poor reputation for consumer sites like this one, so a job that runs on a laptop often fails once it moves to a server.
- CSS class names on result cards change often. Selectors built on them break silently; the JSON keys inside __NEXT_DATA__ change less.
- Third-party tools report different bot vendors for Realtor.com, and none is confirmed by the site. Plan for fingerprint and rate checks in general rather than tuning against one product.
A setup that holds up.
The proxy handles IP reputation and how much traffic each address carries. The browser, pacing, and how you split searches handle the rest.
- Check the free data first: If you need market trends rather than listings, Realtor.com publishes monthly housing inventory metrics (active listings, median list price, days on market) that the St. Louis Fed republishes in FRED. Scrape only when you need listing-level records.
- Load search pages in a real browser: Open search URLs in Playwright and read the __NEXT_DATA__ script once the page has loaded. You get prices, addresses, beds, baths, and permalinks as JSON instead of parsing result cards.
- Split markets into small searches: Work city by city or ZIP by ZIP and walk the pg-<n> pages, instead of paging deep into one huge search. Smaller searches finish faster and spread traffic evenly.
- Give each browser a static ISP IP: Assign one dedicated US residential ISP IP to each browser context and keep it for the whole session, so cookies and the address stay consistent. Add IPs to add throughput instead of pushing one address harder.
- Pace per IP and back off on 429: Leave several seconds between pages on each IP and keep concurrency low. On a 429 or a challenge page, pause that IP for a while. If it stays blocked, replace it through the management API.
import json
import os
from playwright.sync_api import sync_playwright
# Your assigned Stat proxy, kept in environment variables.
PROXY = {
"server": os.environ["STAT_PROXY_SERVER"], # http://<proxy-ip>:<port>
"username": os.environ["STAT_PROXY_USERNAME"],
"password": os.environ["STAT_PROXY_PASSWORD"],
}
def search_page(url: str) -> dict:
with sync_playwright() as p:
browser = p.chromium.launch(proxy=PROXY)
page = browser.new_page()
response = page.goto(url, wait_until="load")
if response is None or response.status != 200:
status = response.status if response else "no response"
browser.close()
raise RuntimeError(f"Got {status}: pause this IP")
raw = page.locator("script#__NEXT_DATA__").text_content(timeout=15_000)
browser.close()
return json.loads(raw)
data = search_page(
"https://www.realtor.com/realestateandhomes-search/<City>_<ST>/pg-1"
)
props = data.get("props", {}).get("pageProps", {})
print(list(props.keys()))Credentials go in separate fields; Playwright does not accept them embedded in the server URL. The listings usually sit under props.pageProps, but the exact key has moved before, so print the keys and check one page by hand before you build a parser on it.
How many IPs to start with.
Starting points, not limits. Your real number depends on how many pages you load, how often, and whether you also open every listing page.
- A few cities or ZIP codes, daily refresh
- 25 IPs (the plan minimum)
- One or two metros, plus listing pages
- 50 to 100 IPs
- Many metros or several refreshes a day
- 200+ IPs or a private subnet
When this is the wrong fit
Stat IPs are US addresses in a few fixed locations and they do not rotate. That suits Realtor.com, which shows the same listings to any US visitor, but if your pipeline expects a fresh IP on every request, or you plan to run many logged-in Realtor.com accounts, this is the wrong setup. For market-level trends, the free inventory data is cheaper than any scraper.
Static ISP Proxies · Pricing · Contact
Frequently asked questions
Does Realtor.com have a public API?
Not a self-serve one for listings. There is no documented listing API with a public key. Market-level inventory data is published for free, and the site itself loads listings through internal endpoints that are undocumented and can change without notice.
Why does Realtor.com return 429 "This is taking longer than usual"?
That page is a soft rate limit. It shows up when one address loads too many pages too quickly, and published tests report it even through residential IPs with a real browser when the pace is too high. Slow down per IP, cut concurrency, use a real browser, and spread the work over more residential IPs.
Do I need rotating proxies to scrape Realtor.com?
No. A new IP on every request throws away the cookies a browser session builds up and makes each request look like a stranger. A pool of static residential ISP IPs, one per browser and each kept at a modest pace, spreads the load and keeps sessions consistent.
How much does it cost to scrape Realtor.com with Stat Proxies?
Static ISP proxies are $2.50 per IP per month with unlimited bandwidth and a 25 IP minimum, so the smallest plan is $62.50 a month. Listing pages load a lot of photos and scripts, and since bandwidth is not metered, that does not change the bill.
Is it allowed to scrape Realtor.com?
Realtor.com's terms of use restrict automated access, and much listing data comes from MLS feeds with their own licensing rules. Stat Proxies provides the network connection; read Realtor.com's terms and get legal advice before collecting data for a commercial product.