How to scrape Zillow listings without getting blocked
Zillow blocks most automated traffic at the edge with HUMAN Security (formerly PerimeterX), and data centre IP ranges are one of the first things it scores. To collect listing data reliably, load pages in a real browser such as Playwright, route each browser through a static US residential ISP proxy, and keep each IP at a human pace. Stat Proxies static ISP IPs cost $2.50 per IP per month with unlimited bandwidth and a 25 IP minimum.
At a glance
- Recommended
- Static ISP Proxies
- Blocker
- HUMAN Security (PerimeterX)
- Client
- Real browser (Playwright, Puppeteer)
- Data location
- __NEXT_DATA__ JSON in the page
- Bandwidth
- Unmetered on every plan
- Starting at
- $2.50 / proxy / mo
Why Zillow scrapers break.
Zillow protects ordinary listing pages, not just its login or search API, and the block is decided at the CDN before your request reaches Zillow's servers.
- Blocked requests come back as a 403 with an x-px-blocked: 1 header, generated by a CloudFront edge function. Retrying the same request from the same client gets the same answer.
- Data centre IP ranges score badly before any other signal is checked, so cloud servers and cheap datacenter proxies fail fast.
- HUMAN expects its client-side sensor to run and set a _px3 cookie. Plain HTTP clients like requests or curl never run it, so they rarely get past the first page.
- Rotating to a new IP on every request breaks the cookie and session the sensor just established, which looks less like a person, not more.
- Listing pages are heavy. Photos, scripts and JSON add up quickly when you refresh thousands of listings a day on a per-GB proxy plan.
A setup that holds up.
The proxy fixes the IP reputation part of the score. The browser and your pacing handle the rest.
- Use a real browser: Run Playwright or Puppeteer so HUMAN's sensor executes and the session gets a valid cookie. Keep one browser context per proxy IP.
- Give each worker a static ISP IP: Assign one dedicated US residential ISP IP to each browser context and keep it for the life of the session. The address and the cookie stay consistent from page to page.
- Read the embedded JSON: Listing pages ship their data in a script tag with the id __NEXT_DATA__. Parsing that JSON is more stable than DOM selectors and includes fields that are not shown on screen.
- Pace per IP and back off on 403: Spread listings across your IPs so no single address requests faster than a person browsing would. On a 403 with x-px-blocked, pause that IP rather than hammering it, and replace a burned IP through the management API if it stays blocked.
import json
import os
from playwright.sync_api import sync_playwright
# Your assigned Stat proxy, kept in environment variables.
PROXY = {
"server": os.environ["STAT_PROXY_SERVER"], # http://<proxy-ip>:<port>
"username": os.environ["STAT_PROXY_USERNAME"],
"password": os.environ["STAT_PROXY_PASSWORD"],
}
def listing_data(url: str) -> dict:
with sync_playwright() as p:
browser = p.chromium.launch(proxy=PROXY)
page = browser.new_page()
response = page.goto(url, wait_until="domcontentloaded")
if response and response.status == 403:
raise RuntimeError("Blocked: pause this IP before retrying")
raw = page.locator("script#__NEXT_DATA__").text_content()
browser.close()
return json.loads(raw)
data = listing_data("https://www.zillow.com/homedetails/<listing-id>_zpid/")
print(list(data["props"]["pageProps"].keys()))Keep credentials in separate fields as shown; Playwright does not accept them embedded in the server URL. Zillow changes its page structure from time to time, so log the keys you depend on and alert when they disappear.
How many IPs to start with.
Starting points, not limits. Your real number depends on how often you refresh and how much of each page you load.
- One metro, daily refresh
- 25 IPs (the plan minimum)
- Several metros or new-listing alerts
- 50 to 100 IPs
- National coverage
- 200+ IPs or a private subnet
When this is the wrong fit
Every Stat Proxies address is static, dedicated, US-based, and reached over HTTP, HTTPS, or SOCKS5 with a username and password. If this job needs rotating exits or IPs outside the US, this is the wrong network for it. Ask us before you buy and we will tell you straight.
Static ISP Proxies · Pricing · Contact
Frequently asked questions
Why does Zillow return a 403 to my scraper?
Zillow uses HUMAN Security at the CloudFront edge. A 403 with an x-px-blocked: 1 header means the request scored as automated, usually because it came from a data centre IP, never ran HUMAN's browser sensor, or arrived too fast. Retrying the same request will not change the score.
Do I need rotating proxies to scrape Zillow?
No. Rotating to a new IP on every request breaks the session cookie HUMAN sets, which tends to make blocks more likely. A small pool of static residential ISP IPs, one per browser session, keeps each session consistent while you spread the total load across addresses.
Can I scrape Zillow with Python requests instead of a browser?
Rarely for long. Plain HTTP clients do not execute HUMAN's client-side sensor, so they never get the cookie the site expects. A real browser such as Playwright, routed through a residential IP, is the dependable approach.
How much does it cost to scrape Zillow with Stat Proxies?
Static ISP proxies are $2.50 per IP per month with unlimited bandwidth and a 25 IP minimum, so the smallest plan is $62.50 a month. Because bandwidth is not metered, loading full listing pages and photos does not change the bill.
Is it allowed to scrape Zillow?
Zillow's terms of use restrict automated access, and rules on collecting public data vary by use and jurisdiction. Stat Proxies provides the network connection; read Zillow's terms and get legal advice before running collection for a commercial product.