How to scrape Redfin listings without hitting 403s and CAPTCHAs
Redfin has no public listing API, but its map search loads results from an internal JSON endpoint (/stingray/api/gis) that is easier to parse than the HTML. The catch is that Redfin starts returning 403s and CAPTCHAs quickly to cloud IPs and to any single address making repeated requests. Load search pages in a real browser, capture that JSON, and give each browser its own static US residential ISP proxy at a human pace. Stat Proxies static ISP IPs cost $2.50 per IP per month with unlimited bandwidth and a 25 IP minimum.
At a glance
- Recommended
- Static ISP Proxies
- Blocker
- 403s and CAPTCHAs after repeated requests per IP
- Client
- Real browser (Playwright) to capture the JSON
- Data location
- /stingray/api/gis JSON, prefixed with {}&&
- Built-in export
- Download All CSV, capped at 350 homes
- Bandwidth
- Unmetered on every plan
- Starting at
- $2.50 / proxy / mo
Why Redfin scrapers break.
Redfin's data is easy to read once you have it. Getting it repeatedly, from the same machine, across many searches, is the hard part.
- Repeated requests from one IP get cut off fast. Scrapers report 403 Forbidden responses and CAPTCHA pages after a modest number of searches, well before you have covered a metro.
- Data centre and cloud IP ranges are flagged before volume even matters, so a scraper that works on your laptop often fails the day you move it to a server.
- Listing detail pages fill in price, beds and baths with JavaScript after the first response. A plain HTTP fetch can come back with a thin page and empty fields that look like real data.
- Every stingray response starts with the characters {}&& before the JSON. Parsers that skip this step throw errors that look like a block but are not.
- The Download All CSV on search results stops at 350 homes and is switched off in some MLS regions, so it does not scale to a full market.
- Redfin changes its internal endpoints from time to time. Hard-coded query strings break silently; capturing the request the page itself makes holds up longer.
A setup that holds up.
The proxy fixes the IP reputation and per-address volume part. The browser, how you split searches, and your pacing handle the rest.
- Check the free data first: If you need market trends (median sale price, inventory, days on market by metro and smaller areas), Redfin's Data Center publishes them as downloads. Scrape only when you need listing-level records.
- Capture the map search JSON in a browser: Open a Redfin search page in Playwright and read the /stingray/api/gis response the page requests for its map. You get structured listing records without guessing query parameters, and the session looks like a normal visit.
- Split big areas into small searches: Each search returns a limited set of homes, so work ZIP by ZIP or neighbourhood by neighbourhood instead of one search for a whole county. Smaller searches also spread requests more evenly.
- Give each browser a static ISP IP: Assign one dedicated US residential ISP IP per browser context and keep it for the whole session, so cookies and the address stay consistent. Add IPs to add throughput rather than pushing one address harder.
- Pace per IP and back off on 403: Leave seconds, not milliseconds, between searches on each IP. On a 403 or a CAPTCHA page, pause that IP, and if it stays blocked replace it through the management API.
import json
import os
from playwright.sync_api import sync_playwright
# Your assigned Stat proxy, kept in environment variables.
PROXY = {
"server": os.environ["STAT_PROXY_SERVER"], # http://<proxy-ip>:<port>
"username": os.environ["STAT_PROXY_USERNAME"],
"password": os.environ["STAT_PROXY_PASSWORD"],
}
def is_gis(response) -> bool:
return "/stingray/api/gis?" in response.url
def search_results(url: str) -> dict:
with sync_playwright() as p:
browser = p.chromium.launch(proxy=PROXY)
page = browser.new_page()
with page.expect_response(is_gis, timeout=30_000) as info:
page.goto(url, wait_until="domcontentloaded")
response = info.value
if response.status != 200:
raise RuntimeError(f"Got {response.status}: pause this IP")
body = response.text()
browser.close()
# Redfin prefixes stingray JSON with {}&& to block script inclusion.
return json.loads(body.removeprefix("{}&&"))
data = search_results("https://www.redfin.com/zipcode/<zip-code>")
print(list(data.get("payload", {}).keys()))Credentials go in separate fields; Playwright does not accept them embedded in the server URL. If expect_response times out, the page may have changed which request feeds the map: open the search in a normal browser, check the Network tab for stingray calls, and update is_gis.
How many IPs to start with.
Starting points, not limits. Your real number depends on how many searches you run, how often, and whether you also open listing pages.
- A few ZIP codes, daily refresh
- 25 IPs (the plan minimum)
- One or two metros, plus listing pages
- 50 to 100 IPs
- Many metros or near-real-time new listings
- 200+ IPs or a private subnet
When this is the wrong fit
Stat IPs are US addresses in a few fixed locations and they do not rotate. That is fine for Redfin, which serves the same listings to any US visitor, but if your pipeline is built around a fresh IP on every request, or you need to log in to many Redfin accounts, this is the wrong setup. For aggregate market trends, Redfin's own Data Center downloads are cheaper than any scraper.
Static ISP Proxies · Pricing · Contact
Frequently asked questions
Does Redfin have a public API?
Not for listings. There is no self-serve key or documented listing API. Redfin's Data Center offers downloadable market-level data, and the site itself loads listings from internal stingray endpoints, which are undocumented and can change without notice.
Why does my Redfin scraper get 403 Forbidden?
Usually because too many requests came from one address, or the address belongs to a cloud or data centre range. Retrying from the same IP rarely helps. Spread searches over more residential IPs, slow down per IP, and use a real browser so the session looks like a normal visit.
Can I just use the Download All CSV?
For small jobs, yes. Redfin caps the download at 350 homes per search, requires at least 20 results, and disables it in some MLS regions. It works for a handful of neighbourhoods but not for covering a full market on a schedule.
Do I need rotating proxies to scrape Redfin?
No. A new IP on every request throws away the cookies a browser session builds up. A pool of static residential ISP IPs, one per browser and each kept at a modest pace, spreads the load without making every request look like a stranger.
How much does it cost to scrape Redfin with Stat Proxies?
Static ISP proxies are $2.50 per IP per month with unlimited bandwidth and a 25 IP minimum, so the smallest plan is $62.50 a month. Listing pages and photos are heavy, and since bandwidth is not metered, loading them does not change the bill.
Is it allowed to scrape Redfin?
Redfin's terms of use restrict automated access, and listing data also comes with MLS rules that vary by region. Stat Proxies provides the network connection; read Redfin's terms and get legal advice before collecting data for a commercial product.