How to Scrape Google People Also Ask Questions at Scale (Python)
To scrape People Also Ask questions at scale, send each seed keyword to an API that renders Google's results page, keep the questions from the People Also Ask box, then dedupe, group and export them. This tutorial builds that pipeline in Python on Searlo's People Also Ask API, at 3 credits per keyword.
What you will build
A script that takes a file of seed keywords and writes two CSV files a writer can use straight away:
• paa_questions.csv: one row per unique question, with its search intent, its question type, how many of your seeds surfaced it and the other ways Google phrased it.
• content_briefs.csv: one row per seed keyword, with its People Also Ask questions in the order Google shows them (a ready FAQ section) and its related searches (secondary topics for the brief).
The steps are: fetch the results page for every seed with POST /serp/batch, read peopleAlsoAsk and relatedSearches from each row, normalise and dedupe the questions, label them by intent, then write the CSVs. The whole script is about 150 lines of Python and needs only requests.
Why not parse Google yourself
The People Also Ask box only exists on a rendered Google results page. Doing it yourself means running browsers and proxies, and keeping your selectors working every time the page changes. Google's own Custom Search JSON API has no People Also Ask field in its response, and Google's overview page says the service will be discontinued on January 1, 2027 (both read 2026-09-29; our migration guide covers the move). Searlo renders the page and returns the box as JSON, and the People Also Ask API page documents the full response.
What comes back for each keyword
Two keys matter for this job:
• peopleAlsoAsk: an array of { "question": "..." }, in the order Google shows them. It holds the question text only. Google loads each answer when someone clicks a question, and Searlo does not click.
• relatedSearches: an array of { "query": "..." }, the related searches Google lists at the foot of the page.
The same row also carries organic, ads, videos and knowledgeGraph. One detail if you use the organic results too: on /serp a result's address is url, not link as on /search/web. Here is a trimmed batch response. The keys are real and the values are illustrative:
json{ "schema": "serp-batch.v1", "results": [ { "keyword": "how to descale a kettle", "success": true, "searchParameters": { "q": "how to descale a kettle", "gl": "us", "hl": "en", "page": 1, "pages": 1, "num": 10, "device": "desktop", "type": "serp" }, "organic": [ { "position": 1, "title": "How to descale a kettle", "url": "https://www.example.com/descale-kettle", "domain": "example.com", "snippet": "…", "page": 1 } ], "peopleAlsoAsk": [ { "question": "How do I descale a kettle?" }, { "question": "Can you descale a kettle with vinegar?" }, { "question": "How often should you descale a kettle?" } ], "relatedSearches": [ { "query": "how to descale a kettle with lemon" }, { "query": "kettle descaler" } ], "fetchedAt": "2026-09-29T09:14:03.512Z", "cached": false }, { "keyword": "kettle limescale", "success": false, "status": 429, "retryable": true, "message": "Google SERP is at capacity. Retry in 60s — immediate retries will be refused." } ], "credits": { "used": 6 }, "meta": { "requested": 2, "distinct": 2, "succeeded": 1, "failed": 1, "pages": 1, "gl": "us", "hl": "en", "device": "desktop", "max_sync_batch": 20 } }
A row that failed carries success: false with a status, a retryable flag and a message in place of the page. The script below handles both kinds of row.
Step 1: settings and one request helper
Put your key in the SEARLO_API_KEY environment variable (a new account comes with 3,000 free credits). Every request goes through one helper, call(). When Searlo is busy it answers 429 or 503 with a Retry-After header, and the helper waits exactly that long before trying again, up to three times. Retrying sooner does not help: immediate retries are refused. Any other error, including a 503 without the header, is raised straight away. The timeout is generous because each page is rendered in a real browser.
python# paa_pipeline.py, step 1: settings and one request helper import csv import os import re import time import unicodedata import requests API = "https://api.searlo.tech/api/v1" HEADERS = {"x-api-key": os.environ["SEARLO_API_KEY"]} MARKET = {"gl": "us", "hl": "en"} def call(method, path, **kwargs): """One request. Waits and retries only when Searlo says so: a 429 or 503 with Retry-After.""" for attempt in range(4): resp = requests.request(method, API + path, headers=HEADERS, timeout=600, **kwargs) wait = resp.headers.get("retry-after") if resp.status_code in (429, 503) and wait and attempt < 3: time.sleep(int(wait)) continue resp.raise_for_status() return resp.json()
Step 2: check one keyword first
Before you spend credits on a list, look at one page. GET /serp with aioverview=false skips waiting for the AI Overview, which question mining does not need. The price is 3 credits either way.
python# paa_pipeline.py, step 2: check one keyword by hand before you batch anything def one_serp(keyword): return call("GET", "/serp", params={"q": keyword, **MARKET, "aioverview": "false"}) # In a Python shell: # >>> serp = one_serp("how to descale a kettle") # >>> [item["question"] for item in serp["peopleAlsoAsk"]] # >>> [item["query"] for item in serp["relatedSearches"]]
If the questions look right for your market, run the list.
Step 3: fetch every seed, 20 per request
POST /serp/batch takes up to 20 keywords and costs 3 credits per distinct keyword per results page. The script lowercases and dedupes the seeds first, then sends them 20 at a time with the same gl and hl. The AI Overview is off by default on the batch endpoint, so there is nothing to switch off.
One billing rule shapes the retry logic. A row that fails inside a batch is still billed, because the batch request as a whole succeeded. A single GET /serp that fails is refunded automatically. So retry_failed() re-runs the failed keywords one at a time: if a keyword fails again, it costs nothing.
python# paa_pipeline.py, step 3: every seed keyword, 20 per POST /serp/batch def fetch_serps(seeds): keywords = list(dict.fromkeys(s.strip().lower() for s in seeds if s.strip())) ok, failed = [], [] for i in range(0, len(keywords), 20): batch = call("POST", "/serp/batch", json={"keywords": keywords[i:i + 20], **MARKET}) for row in batch["results"]: (ok if row["success"] else failed).append(row) meta = batch["meta"] print(f"{meta['succeeded']} ok, {meta['failed']} failed, {batch['credits']['used']} credits") return ok, failed def retry_failed(failed): """Re-run failed rows one at a time: a failed single GET /serp is refunded.""" rows = [] for row in failed: try: rows.append({"keyword": row["keyword"], **one_serp(row["keyword"])}) except requests.HTTPError as err: print("still failing:", row["keyword"], err) return rows
Step 4: collect and dedupe the questions
The same question turns up under several seeds, often phrased a little differently ("How do I descale a kettle?", "How to descale a kettle"). The script builds a key for each question: Unicode-normalised, lowercased, punctuation removed, a handful of filler words dropped, and the remaining words sorted. Questions with the same key merge into one record. The record keeps the first phrasing, every other phrasing, the seeds that surfaced it and its best position in the box.
The filler-word list is English. Adapt it for other languages, and keep it short: the aim is to merge rewordings, not to merge different questions.
python# paa_pipeline.py, step 4: one record per question, remembering which seeds surfaced it STOPWORDS = {"a", "an", "the", "i", "you", "we", "my", "your", "our", "do", "does", "to"} def normalise(text): text = unicodedata.normalize("NFKC", text).casefold() text = re.sub(r"[^\w\s]", " ", text) return re.sub(r"\s+", " ", text).strip() def signature(text): """Order-free key, so 'How do I descale a kettle?' and 'How to descale a kettle' merge.""" return " ".join(sorted(set(normalise(text).split()) - STOPWORDS)) def collect(rows): questions, briefs = {}, [] for row in rows: seed = row["keyword"] paa = [item["question"] for item in row["peopleAlsoAsk"]] related = [item["query"] for item in row["relatedSearches"]] briefs.append({"seed": seed, "questions": paa, "related": related}) for position, question in enumerate(paa, start=1): rec = questions.setdefault(signature(question), { "question": question, "variants": set(), "seeds": set(), "best_position": position, }) rec["variants"].add(question) rec["seeds"].add(seed) rec["best_position"] = min(rec["best_position"], position) return questions, briefs
Step 5: group the questions by intent
Each question gets two labels.
• Search intent, from GET /keywords/intent: informational, navigational, commercial or transactional. It costs 0 credits. It classifies from the words alone (its response says "basis": "lexical"), so treat it as a first sort, not a verdict. The script sends 50 questions per call to keep the URL short, strips the trailing question mark, and replaces commas with spaces because a single keyword string is split on commas.
• Question type, from the question's opening words: cost, time, how-to, comparison, definition, reason or yes/no. It tells the writer what kind of answer the question expects. Order matters in the list: "how much" is tested before "how", and "best way" counts as how-to before "best" counts as comparison.
Together they route the work. Informational how-to and definition questions become FAQ entries and guide sections. Commercial comparison questions belong on comparison pages. Transactional cost questions belong on pricing and product pages.
python# paa_pipeline.py, step 5: group by intent: Searlo's free intent label, then the question's shape QUESTION_TYPES = [ ("cost", r"^how much\b|\b(cost|costs|price|prices)\b"), ("time", r"^(when|how (long|often|soon))\b"), ("how-to", r"^how\b|\bbest way\b"), ("comparison", r"\b(vs|versus|better|best|difference|compare)\b"), ("definition", r"^(what|who)\b"), ("reason", r"^why\b"), ("yes/no", r"^(can|could|does|do|is|are|should|will|would)\b"), ] def question_type(question): q = normalise(question) return next((label for label, pattern in QUESTION_TYPES if re.search(pattern, q)), "other") def search_intent(questions): """GET /keywords/intent: lexical labels, 0 credits. 50 per call keeps the URL short.""" labels = {} for i in range(0, len(questions), 50): chunk = questions[i:i + 50] sent = [q.rstrip("?").replace(",", " ") for q in chunk] data = call("GET", "/keywords/intent", params={"keywords": sent}) for question, result in zip(chunk, data["results"]): labels[question] = result["intent"] return labels
Step 6: export CSVs for briefs and FAQ sections
paa_questions.csv is sorted by search intent, then question type, then by how many seeds surfaced the question. A question that appears under many of your seeds is a core question for the topic, and a strong candidate for the FAQ of your main page. content_briefs.csv has one row per seed. Its questions sit in one cell, one per line, in the order Google showed them.
python# paa_pipeline.py, step 6: two CSVs, one row per question and one brief per seed keyword def export(questions, briefs, intents): records = sorted(questions.values(), key=lambda r: ( intents.get(r["question"], ""), question_type(r["question"]), -len(r["seeds"]), r["best_position"], )) with open("paa_questions.csv", "w", newline="", encoding="utf-8") as f: w = csv.writer(f) w.writerow(["search_intent", "question_type", "question", "seed_count", "seeds", "best_position", "other_phrasings", "gl", "hl"]) for r in records: w.writerow([ intents.get(r["question"], ""), question_type(r["question"]), r["question"], len(r["seeds"]), "; ".join(sorted(r["seeds"])), r["best_position"], "; ".join(sorted(r["variants"] - {r["question"]})), MARKET["gl"], MARKET["hl"], ]) with open("content_briefs.csv", "w", newline="", encoding="utf-8") as f: w = csv.writer(f) w.writerow(["seed_keyword", "faq_questions", "related_searches", "paa_found"]) for b in briefs: w.writerow([b["seed"], "\n".join(b["questions"]), "; ".join(b["related"]), len(b["questions"])])
The first rows of paa_questions.csv look like this (illustrative values):
csvsearch_intent,question_type,question,seed_count,seeds,best_position,other_phrasings,gl,hl commercial,comparison,"Which is better for descaling, vinegar or lemon?",1,vinegar vs lemon descaling,1,,us,en informational,how-to,How do I descale a kettle?,2,how to descale a kettle; kettle limescale,1,How to descale a kettle,us,en informational,time,How often should you descale a kettle?,1,how to descale a kettle,3,,us,en informational,yes/no,Is limescale in a kettle harmful?,2,how to descale a kettle; kettle limescale,3,,us,en
Step 7: run it
Put one keyword per line in seeds.txt, set SEARLO_API_KEY, and run python paa_pipeline.py. The batch step prints the credits each batch used. To cover another market, change MARKET (for example {"gl": "gb", "hl": "en"}) and run it again into a separate folder, because the questions differ by country.
python# paa_pipeline.py, step 7: run it on a seeds.txt with one keyword per line if __name__ == "__main__": with open("seeds.txt", encoding="utf-8") as f: seeds = [line for line in f if line.strip()] ok, failed = fetch_serps(seeds) ok += retry_failed(failed) questions, briefs = collect(ok) intents = search_intent([r["question"] for r in questions.values()]) export(questions, briefs, intents) print(len(questions), "unique questions from", len(briefs), "seed keywords")
Optional: go one level deeper
Every question you found is a keyword too, and searching it returns its own People Also Ask box. That is how a list of 50 seeds turns into a question map for a whole topic. Each expanded question is another results page at 3 credits, so cap the expansion. The snippet below searches the questions that the most seeds share. Run it after collect(), then carry on with search_intent() and export() as before. The seeds column then also records which question surfaced each new one.
python# Optional, after step 4: expand one level by searching the questions themselves (3 credits each) MAX_EXPAND = 100 top = sorted(questions.values(), key=lambda r: (-len(r["seeds"]), r["best_position"]))[:MAX_EXPAND] more_ok, more_failed = fetch_serps([r["question"] for r in top]) more_ok += retry_failed(more_failed) questions, briefs = collect(ok + more_ok)
To widen the seed list before you start, the Google Autocomplete API returns the suggestions Google offers as a query is typed, at 0.5 credits per call.
What it costs
The formula is 3 credits × distinct keywords × markets × runs. The intent labels are free. Keep the default of one results page: People Also Ask and related searches come from page one, and each extra page adds organic results only, for another 3 credits per keyword. The dollar figures below use each pack's published rate from our pricing page.
A new account's 3,000 free credits cover the first 1,000 results pages, which is enough to run the first two jobs above. Packs are one-time purchases, not a subscription, and credits on new accounts are valid for 90 days, so pick the pack you will use inside that window: the smallest is 5,000 credits for $3.99, and Starter is 20,000 credits for $9.99.
Honest limits
• Not every results page has a People Also Ask box. Many commercial and navigational queries show none. peopleAlsoAsk is then an empty array and the page is still billed, because it was fetched. Those seeds show paa_found = 0 in content_briefs.csv.
• The box changes with country, language, device and time. The same keyword can surface different questions with gl=us and gl=gb, on device=mobile and desktop, and from one month to the next. Keep the market and the run date with the data, and run each market on its own.
• Questions only, and only the first set. An item is { question }. Google loads the answers, and more questions, when someone clicks; Searlo does not click, so you get what a first-time visitor sees. To see who answers a question, search the question itself with GET /search/web (1 credit).
• Page one only. pages=2 or more adds organic results, not questions.
• Question-shaped entries only. The parser keeps entries that end in a question mark or start with an English question word (how, what, why, can, is and so on). In languages that end questions differently, some entries can be missed, so check a sample before you scale to a new market.
• Failed batch rows are billed. A single GET /serp that fails is refunded, which is why the script re-runs failures one at a time.
• Seconds, not milliseconds. Each page is rendered in a real browser and a batch fetches a few keywords at a time, so a full batch of 20 takes a while. For hundreds of keywords at once, queue async tasks instead; the Bulk SERP API page has a 500-keyword example.
• Cached pages cost the same. A request repeated soon after an identical one can be answered from Searlo's short-lived cache ("cached": true), and it is billed like a fresh page. Send cache: false in the batch body, or cache=false on GET /serp, when you need a new fetch.
Where to go next
• People Also Ask API: the full response, the fields table and more code samples.
• Bulk SERP API: async tasks and webhooks for lists in the hundreds or thousands.
• Google Autocomplete API: seed expansion from Google's suggestions.
• AI Overview API: the AI Overview for the same query, with aioverview=true on the same endpoint.
• API reference and docs: every parameter, including location, uule and device for city-level and mobile results.
• Pricing: every pack and its rate per 1,000 credits.
Searlo is an independent service and is not affiliated with or endorsed by Google. Google is a trademark of Google LLC. Searlo returns Google's publicly visible results page as JSON.
Frequently Asked Questions
How do I scrape Google People Also Ask questions with Python?
Send your keywords, 20 at a time, to POST /serp/batch on Searlo's API, read the question field of each item in peopleAlsoAsk, dedupe the questions and write them to CSV. Each keyword costs 3 credits per results page. The tutorial above has a tested script of about 150 lines that also groups the questions by intent and writes a content brief per keyword.
Can I get the answers to People Also Ask questions?
No. Each peopleAlsoAsk item is the question text only, because Google loads an answer when someone clicks a question and Searlo does not click. To see who answers a question, search the question itself: GET /search/web returns the organic results for 1 credit, and GET /serp returns the whole results page for 3.
How much does it cost to scrape People Also Ask for 1,000 keywords?
1,000 keywords in one market is 3,000 credits: $1.50 at the Starter pack's rate, $0.90 at the Scale rate and $0.66 at the Pro rate. A new account's 3,000 free credits cover it once. Keywords whose results page has no People Also Ask box are billed the same, because the page was fetched.
Why did a keyword return no People Also Ask questions?
Most often because that results page has no People Also Ask box, which is common for commercial and navigational queries. The box also varies by country, language, device and date, and the parser keeps only question-shaped entries, so a market whose questions do not end in a question mark can return fewer. The page is billed either way. Check a few such keywords with GET /serp before you drop them.
Do People Also Ask questions change by country or device?
Yes. Set gl for the country, hl for the language, device for desktop, mobile or tablet, and location for a city. Run each market separately and keep the run date with the data, because the same keyword can surface different questions in different markets and from month to month.
How many keywords can I send in one request?
Up to 20 per POST /serp/batch, and duplicates in a batch are billed once. For hundreds or thousands of keywords, queue async tasks and collect the results later or by webhook; the Bulk SERP API page has a 500-keyword example. Failed async tasks are refunded, while a failed row inside a synchronous batch is billed.
Is Searlo affiliated with Google?
No. Searlo is an independent service and is not affiliated with or endorsed by Google. Google is a trademark of Google LLC. Searlo returns Google's publicly visible results page as JSON.