How to Solve CAPTCHA in Crawl4AI Without Losing the Session

Crawl4AI has no CAPTCHA solver of its own, so a Crawl4AI captcha flow is three calls you wire together: load the page in a named session, solve the sitekey with an external solver, then go back into the same tab to put the token in and submit. The part that catches people is the order Crawl4AI runs things in. Since version 0.8.5, js_code runs after wait_for, so a crawl that submits the form in js_code and waits for the next page with wait_for sits there until it times out. The submit belongs in js_code_before_wait. This guide walks through the session flow with reCAPTCHA v2, then a hook for deep crawls where you do not control each call.
What you need
- Python 3.10 or newer and Crawl4AI 0.9 or later. Everything here was read from the 0.9.4 source and run against it; the ordering in Step 3 goes back to 0.8.5.
- The CapSkip Python package, which gives you AsyncCapSkip, a genuinely asynchronous client that fits the event loop Crawl4AI already runs on.
- The URL of the page with the CAPTCHA, and a selector for something that only appears once the form has gone through.
- CapSkip running on a Windows machine. Local mode answers on 127.0.0.1 when the crawler runs on the same box; Server mode listens on your network or public IP when it does not. Both are covered under connection settings.
# pip install crawl4ai pip install -U crawl4ai capskip # Downloads the browser Crawl4AI drives, once per machine. crawl4ai-setup
Step 1: load the page in a session and read the sitekey
Give the first call a session_id. That keeps the tab open after the call returns, with its cookies and its widget intact, so the token you solve later lands in the page that asked for it. Without one, Crawl4AI closes the page as soon as it has the HTML.
# pip install crawl4ai
import re
from crawl4ai import CrawlerRunConfig
PAGE_URL = "https://example.com/signup"
SESSION = "signup"
# The widget element, in any attribute order, with or without other classes.
WIDGET = r'<[^>]*class="(?:[^"]*\s)?g-recaptcha(?:\s[^"]*)?"[^>]*>'
async def read_sitekey(crawler):
# session_id keeps this tab open for the next arun call.
config = CrawlerRunConfig(session_id=SESSION)
first = await crawler.arun(PAGE_URL, config=config)
# Do not stop on first.success: a CAPTCHA page can be
# flagged as blocked while its HTML is complete.
widget = re.search(WIDGET, first.html)
match = widget and re.search(r'data-sitekey="([^"]+)"', widget.group(0))
return match.group(1) if match else NoneThat comment is there for a reason. Crawl4AI 0.9 runs an anti-bot check on every result, and a page it reads as a block page is marked failed: success comes back False and error_message starts with Blocked by anti-bot protection. Any 403 or 503 HTML response qualifies, and so does a 429. On other error statuses, a page under 10 KB qualifies when its widget’s class attribute is exactly g-recaptcha. The rendered HTML is still in first.html, and it still holds the sitekey you need, so read it before you decide the crawl failed.
The html field is the rendered page, so a widget built from script is in it too. If the key is not on the element, look for the k= parameter in the src of the widget iframe, and add a wait_for on that iframe to the first call if the widget renders late.
Step 2: solve it with AsyncCapSkip
Crawl4AI is asyncio from top to bottom, so use the asynchronous client. It is a real coroutine rather than a wrapper around blocking calls, which means a solve that takes twenty seconds does not freeze every other crawl on the same event loop.
# pip install capskip
from capskip import AsyncCapSkip
solver = AsyncCapSkip(host="127.0.0.1", port=8080)
async def solve(sitekey):
# Invisible v2 takes invisible=1, Enterprise takes enterprise=1.
result = await solver.recaptcha(sitekey=sitekey, url=PAGE_URL)
return result["code"] # the g-recaptcha-response valueNothing in Crawl4AI is waiting while this runs, because the solve happens between two calls, not inside one. The tab simply sits there. What does have a clock is the token: a reCAPTCHA token is good for about two minutes after it is issued, so go straight on to the submit. The details are in how long a reCAPTCHA token stays valid.
Step 3: inject and submit in the same tab
The second call reuses the session and sets js_only, which tells Crawl4AI to run JavaScript in the page it already has instead of loading the URL again. A reload costs a second page load, which can be challenged all over again, and it resets anything the first load set up, such as a partly filled form.
import json
async def submit(crawler, token):
# Crawl4AI wraps this in an async function, so statements work.
inject = (
"document.getElementById('g-recaptcha-response').value = "
f"{json.dumps(token)};"
"document.querySelector('form').submit();"
)
config = CrawlerRunConfig(
session_id=SESSION,
js_only=True, # same tab, no reload
js_code_before_wait=inject, # runs BEFORE wait_for
wait_for="css:.signup-complete",
)
return await crawler.arun(PAGE_URL, config=config)Here is why the script goes in js_code_before_wait. Since 0.8.5, and in every 0.9 release, a crawl runs js_code_before_wait first, then wait_for, then js_code last, against the finished page. So if you put the submit in js_code, wait_for starts looking for the next page before anything has been submitted, and it keeps looking until page_timeout runs out, 60 seconds by default, when the call fails with Wait condition failed and the form is never sent. Examples that submit from js_code and wait with wait_for in the same call, including one in Crawl4AI’s own documentation, run into exactly this on 0.9.4.
Build the string with json.dumps rather than pasting the token between quotes, because a JSON string literal is also a valid JavaScript one, escaping included. You do not have to wait for the navigation yourself: wait_for keeps polling across it and finds .signup-complete on the next page. What Crawl4AI will not do is raise when the script fails. A syntax error gets a log line, but a run-time error, such as an element ID with a typo in it, is swallowed without one and shows up only as the wait_for timing out, so test the two statements in the DevTools console on the real page before you blame the solver.
Some sites never read the textarea. They register a callback with the widget and submit from there, so replace the two statements with a call to the function named in the data-callback attribute, passing the token. It is still ordinary reCAPTCHA v2 underneath, as explained on the reCAPTCHA v2 solver page.
Step 4: solving inside a deep crawl with a hook
The session flow handles a Crawl4AI captcha when you own each arun call. A deep crawl or an arun_many batch does not give you that, so attach the solve to the after_goto hook instead. It runs on the Playwright page right after navigation and before wait_for, for every URL the crawler navigates to (js_only calls skip it), so pages without a widget pass straight through.
from capskip import CapSkipError
async def after_goto(page, context, url, response, **kwargs):
holder = await page.query_selector("div.g-recaptcha[data-sitekey]")
if holder is None:
return page # no widget, nothing to do
sitekey = await holder.get_attribute("data-sitekey")
try:
result = await solver.recaptcha(sitekey=sitekey, url=page.url)
except CapSkipError:
return page # keep the page HTML if the solve fails
async with page.expect_navigation():
await page.evaluate(
"t => { document.getElementById('g-recaptcha-response').value = t;"
" document.querySelector('form').submit(); }", result["code"])
return page
crawler.crawler_strategy.set_hook("after_goto", after_goto)A few things to know about this route. The hook runs inside the crawl, so the solve time is added to that page, and several pages hitting a widget at once each wait on their own solve, which AsyncCapSkip handles side by side. The hook is global to the crawler, so keep the check cheap, because every page runs it. And the try block matters: an exception that escapes the hook fails the whole crawl of that URL and returns no HTML at all, so a CapSkip outage would otherwise empty every protected page in a deep crawl. One more catch: Crawl4AI keeps the status code of the first response. A challenge served as a 403 still comes back with success False and Blocked by anti-bot protection, even when the hook solved it and result.html is the page behind it. On those URLs, check the html for your content rather than trusting success. Crawl4AI documents every hook and its arguments on its hooks page. The same hook suits any Playwright-driven crawler, and the wider picture is on the Playwright CAPTCHA solver page.
Where CapSkip runs when the crawler does not
Crawl4AI often ends up on a Linux server or in a container once a crawl grows, and CapSkip is a Windows application, so the two frequently live on different machines. That is what Server mode is for. It makes CapSkip listen on your network address or public IP instead of 127.0.0.1, so the crawler calls it over the same HTTP API from anywhere it can route to. Use a static public IP if that route crosses the internet, turn on API key validation, and restrict the port to the crawler’s addresses with a Windows Firewall rule. It is still a machine you own, and solving on it is still unmetered.
The SDK does not read environment variables by itself. Read CAPSKIP_HOST, CAPSKIP_PORT and CAPSKIP_API_KEY in your own code and pass them in, as the full example does.
One Crawl4AI detail to plan around: since 0.9.0 its Docker server rejects session_id, js_code and js_code_before_wait with an HTTP 400 when they arrive over the network, and it no longer accepts hook code. Neither route in this guide can go through the REST API, so run it with the library in your own Python process.
Full working example
# pip install crawl4ai capskip
import asyncio
import json
import os
import re
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
from capskip import AsyncCapSkip
PAGE_URL = "https://example.com/signup"
SESSION = "signup"
WIDGET = r'<[^>]*class="(?:[^"]*\s)?g-recaptcha(?:\s[^"]*)?"[^>]*>'
solver = AsyncCapSkip(
apiKey=os.getenv("CAPSKIP_API_KEY", "capskip"),
host=os.getenv("CAPSKIP_HOST", "127.0.0.1"),
port=int(os.getenv("CAPSKIP_PORT", "8080")),
)
async def main():
async with AsyncWebCrawler() as crawler:
first = await crawler.arun(
PAGE_URL, config=CrawlerRunConfig(session_id=SESSION))
widget = re.search(WIDGET, first.html)
match = widget and re.search(r'data-sitekey="([^"]+)"', widget.group(0))
if not match:
raise RuntimeError(f"no sitekey found: {first.error_message}")
result = await solver.recaptcha(sitekey=match.group(1), url=PAGE_URL)
inject = (
"document.getElementById('g-recaptcha-response').value = "
f"{json.dumps(result['code'])};"
"document.querySelector('form').submit();"
)
done = await crawler.arun(PAGE_URL, config=CrawlerRunConfig(
session_id=SESSION,
js_only=True,
js_code_before_wait=inject,
wait_for="css:.signup-complete",
wait_for_timeout=30000,
))
print(done.success) # True once the next page has loaded
asyncio.run(main())wait_for_timeout gives the post-submit wait its own limit, so a submit that goes nowhere fails in 30 seconds rather than in the 60 of page_timeout. Swap .signup-complete for anything that exists only on the page after the form. For the rest of the Python landscape, including Selenium and plain HTTP clients, see the Python CAPTCHA solver page.
Common errors and what they mean
| What you see | Cause | Fix |
|---|---|---|
| Wait condition failed after about 60 seconds, and the form never submitted | The submit is in js_code, which runs after wait_for since 0.8.5 | Move the script to js_code_before_wait |
| first.success is False with Blocked by anti-bot protection | Crawl4AI’s block check read the CAPTCHA page as a block page | Read the sitekey from first.html anyway; the page is intact |
| wait_for times out on the second call and the html is a blank page | The session_id differs, so Crawl4AI opened a new, empty tab | Use the same session_id on both calls |
| The token is in the textarea but the site says the CAPTCHA failed | The site submits through a data-callback function, or the token expired | Call the callback with the token, and submit within two minutes of solving |
| No error at all, then the wait times out | The injection script threw inside the page, and Crawl4AI swallows that silently | Run the script in the DevTools console first and check the IDs it uses |
| HTTP 400 from the Crawl4AI Docker server | The server refuses session_id and script fields that arrive over the network | Run the library in your own Python process |
| NetworkException from the solve | CapSkip is not running, or the host and port point at the wrong machine | Start CapSkip, and use Server mode when the crawler is elsewhere |
| TimeoutException from the solve | The solve outlasted recaptchaTimeout | Raise it above the default of 300 seconds in the constructor |
FAQ
Does Crawl4AI solve CAPTCHAs on its own?
No. It detects them, in the sense that its anti-bot check marks a CAPTCHA page as a blocked crawl, and it can retry through a list of proxies when a page is blocked. The magic and simulate_user options add the mouse movement and scrolling that anti-bot systems look for, which can make a challenge less likely. None of that solves a widget once it is on the page, which is the gap an external solver fills, and it is the same gap described on the CAPTCHA solver for web scraping page.
Does the same flow work for Cloudflare Turnstile?
For the widget, yes. Call solver.turnstile with the sitekey and page URL, and write the token into the hidden input named cf-turnstile-response instead of the reCAPTCHA textarea. Full-page challenges are a different job, because their token has to travel with the user agent CapSkip reports back, so the browser that submits it has to present that same user agent. That case is covered on the Cloudflare Turnstile solver page.
My crawler runs in a container. Where does CapSkip go?
On a Windows machine you control, with Server mode switched on. The container then reaches it over the API like any other internal service, so the crawler and the solver do not need to share an operating system. Pass the Windows machine’s address as CAPSKIP_HOST, enable API key validation, and give the crawler its own key so it can be revoked on its own.
How long does a session stay open between the two calls?
Far longer than a solve takes. The browser manager in the 0.9.4 source cleans up sessions that have gone unused for 30 minutes, and closing the crawler closes them all. The limit that matters in practice is the token’s, about two minutes, so the solve and the submit should follow each other directly.
The short version
Here is the whole Crawl4AI captcha flow in a paragraph. Load the page with a session_id and read the sitekey from the HTML, even when Crawl4AI calls the crawl blocked. Solve it with AsyncCapSkip. Go back into the same tab with js_only=True, put the token in and submit from js_code_before_wait, and wait for the next page with wait_for and its own timeout. For deep crawls, do the same from the after_goto hook. Point the client at a Server mode address when the crawler runs on another machine.
One last thing that changes how you size a crawl. Because the captcha bypass runs on a machine you already own, a crawl that meets a widget on every page costs the same as one that meets it once, so there is no reason to skip a URL just because it is behind a challenge.
