How to Solve ALTCHA in Scrapy With One Inline Request

solve altcha in scrapy - How to Solve ALTCHA in Scrapy With One Inline Request

To solve ALTCHA in Scrapy, read the challenge address off the altcha-widget element, fetch it with an inline request from the same callback, hand the JSON to AsyncCapSkip’s altcha method, and submit the form with the token in the field the widget names, which is altcha by default. ALTCHA is proof of work, not a picture, so none of this needs a browser, and a plain Scrapy spider is enough. The one Scrapy-specific trap is easy to miss: a challenge endpoint returns a new document at the same URL every time, and Scrapy’s duplicate filter quietly drops the second request to it. This guide covers the widget, the fetch, the solve and the submit, with a spider you can run.

What you need

  • Scrapy 2.14 or later, for the inline request method used below. Scrapy has run on the asyncio reactor by default since 2.13, which is what lets a callback await the async client. A section near the end covers older versions.
  • Version 1.2.0 or later of the capskip Python package, the release that added the altcha method, and CapSkip 1.3.0 or later running on a Windows machine.
  • The form2request package, which Scrapy now recommends for building form submissions.
  • The page with the form. Everything else, the challenge included, comes off that page, and Step 1 shows where.
  • An address for the solver. Local mode answers on 127.0.0.1 for that device only; Server mode listens on your network address or public IP so a crawler on another machine can call it over the API. Both live under connection settings, and a section below covers when to switch.
# pip install scrapy capskip form2request
pip install scrapy capskip form2request

Step 1: read the challenge off the widget

ALTCHA has no sitekey. What you need to solve ALTCHA in Scrapy is the challenge, or the address it comes from, and the widget element carries one or the other. Which attribute holds it depends on the widget generation, so read all three.

WidgetAttributeHolds
v1 and v2challengeurl, or challengejsonchallengeurl names the endpoint that serves a challenge, often a relative path; challengejson carries the challenge document itself
v3 and laterchallengeEither that endpoint, or the challenge document itself
# Inside a spider callback; response is the page with the form.
widget = response.css("altcha-widget")
source = (widget.attrib.get("challenge")
          or widget.attrib.get("challengejson")
          or widget.attrib.get("challengeurl"))

# The token goes in a field named by the widget, altcha by default.
field = widget.attrib.get("name", "altcha")

# A v3 challenge or a v1/v2 challengejson carries the document inline.
inline = source.lstrip().startswith("{")

Also note the name attribute. The widget writes its payload into a hidden input with that name, and it defaults to altcha, but a site can change it. Reading it costs one line and saves a submit that silently leaves the real field empty. A few deployments generate the challenge in the page’s own script rather than serving it from an endpoint, and there is then no attribute to read. The Network tab shows which kind you have.

Step 2: fetch the challenge inside the callback

If the widget points at an endpoint, fetch it from the callback you are already in, with the engine’s download_async method, which Scrapy’s documentation describes under inline requests. The request goes through the downloader middlewares, so the spider’s cookie jar, user agent and proxy settings apply to it, as they did to the page. It does not go through the scheduler, and that is the point.

# The page's jar and proxy; no scheduler, no dupefilter.
meta = {"dont_cache": True, "allow_offsite": True}
for key in ("cookiejar", "proxy"):
    if key in response.meta:
        meta[key] = response.meta[key]
reply = await self.crawler.engine.download_async(
    scrapy.Request(response.urljoin(source), meta=meta,
                   headers={"Referer": response.url})
)
if reply.status != 200:
    raise ValueError(f"challenge endpoint answered {reply.status}")
challenge_json = reply.text

Here is why skipping the scheduler matters. The obvious alternative, yielding a Request for the challenge with a callback of its own, works for the first form and loses the second. Every fetch of the endpoint returns a new challenge, but the URL never changes, so Scrapy’s duplicate filter treats the second request as one it has already made and discards it. Nothing errors. The filter logs its first discard once at debug level, bumps a counter in the crawl stats, and the form behind it is simply never submitted. The inline request never reaches the filter.

The rest of that block covers details the scheduler would normally handle for you. Per-request state does not carry over by itself, so the loop copies a cookiejar or proxy key from the page’s meta when the page was fetched with one. The Referer header is set by hand because the middleware that adds it runs on the spider side. If the challenge endpoint sits on another domain, a hosted ALTCHA service for example, and the spider sets allowed_domains, the offsite filter would block the fetch, and allow_offsite lets this one request through. The status check exists because a 403 from the endpoint comes back as an ordinary response here: the middleware that turns error statuses into failures runs on the spider side, which an inline request skips. And dont_cache keeps Scrapy’s HTTP cache away from the challenge if you have it on. A cached challenge is an expired one, and CapSkip refuses an expired inline challenge straight away rather than spend CPU on a token the site will reject. Without the flag, a second run hands every page that shares the endpoint the same stored challenge. The same logic covers the form page. A cached page replays an old anti-forgery token and session cookie, and with an inline challenge an old challenge too, so with the cache on, put dont_cache on the form pages as well, or leave the cache off for this spider.

You could instead pass the endpoint as challenge_url and let CapSkip fetch the challenge itself. That is fine for a public endpoint. But CapSkip’s fetch carries none of your spider’s cookies, and some sites only issue challenges to the session that loaded the form. Fetching through Scrapy keeps one session end to end.

Step 3: solve it without blocking the crawl

Use AsyncCapSkip and make the callback async def. In the Python package that class is a genuine asyncio client rather than an alias, so awaiting it hands control back to Scrapy while the solve runs and every other request keeps moving. The synchronous CapSkip class would stop the whole crawl for the length of each solve, a problem explained in detail in the Scrapy CAPTCHA middleware guide.

from capskip import AsyncCapSkip

solver = AsyncCapSkip(host="127.0.0.1", port=8080)

# Pass the document exactly as the endpoint sent it.
result = await solver.altcha(url=response.url, challenge_json=challenge_json)

token = result["token"]   # base64 payload for the form field

Because the challenge travels inline, CapSkip makes no network request at all. It hashes until it finds the answer, which usually takes milliseconds, and the client polls straight away and again a quarter of a second later, so a typical solve returns in about a quarter of a second. The call waits up to defaultTimeout, 120 seconds, because this is CPU work rather than a browser session. Two challenge algorithms, Argon2id and scrypt, are refused rather than attempted, and that refusal arrives as an ApiException in under a second. The more common PBKDF2 and SHA schemes, including ALTCHA’s recommended default, are covered.

Pass the reply’s text as it arrived; there is no need to parse it. What must not change is the token: it is base64 of a JSON document whose fields the site’s server has signed, so a token that was trimmed, decoded or re-encoded fails verification.

Step 4: submit the form with the token

The hidden input the widget fills does not exist in the HTML Scrapy downloaded. The widget creates it in a browser, after it has run, so nothing that reads the served HTML can find it. Add it yourself, under the name from Step 1. The example builds the submit with form2request, which Scrapy recommends in place of FormRequest.from_response: since Scrapy 2.16 that older method logs a deprecation warning on every call.

from form2request import form2request

form = response.xpath("//form[.//altcha-widget]")
data = {"email": "[email protected]", field: token}
yield form2request(form, data).to_scrapy(
    callback=self.after_submit,
    priority=10,
    meta={"handle_httpstatus_all": True},
)

form2request keeps what the page already put in the form, such as an anti-forgery token, and presses the first submit button, so the request looks like the one a browser would send. The XPath picks the form that contains the widget, which matters on pages that also carry a search box. handle_httpstatus_all lets the callback see a rejected submit, which Scrapy would otherwise drop before it arrives. The priority bump sends the submit ahead of everything still waiting in the scheduler, because challenges expire. Some windows are barely two minutes long, and a token that waits behind a long crawl can expire before it is sent.

Treat each token as good for one submit. If the site rejects the form, start again from Step 2 with a new challenge rather than retrying the same request. Some integrations send the payload in a JSON body or a cookie instead of a form field, so submit the form once by hand with DevTools open and mirror what the page sends.

Running the solver somewhere else

The samples use 127.0.0.1 because that is right while the spider and CapSkip share a machine. A crawler deployed to Scrapyd on another box, a VPS or a hosted Scrapy platform points loopback at itself, and the first solve raises a NetworkException. Switch CapSkip to Server mode and it listens on your network address or public IP, so the spider can reach it over the API from anywhere you allow. Use a static public IP if the route crosses the internet, turn on API key validation, and restrict the port to the addresses you expect with a Windows Firewall rule. It is still your own Windows machine, and solving is still unmetered.

The client does not read environment variables by itself. Read CAPSKIP_HOST and CAPSKIP_API_KEY in the spider and pass them to the constructor, as the full example does, so the same code runs on your desk and on the server.

Full working example

# pip install scrapy capskip form2request
import os

import scrapy
from capskip import AsyncCapSkip, CapSkipError
from form2request import form2request


class SignupSpider(scrapy.Spider):
    name = "signup"
    start_urls = ["https://example.com/signup"]

    solver = AsyncCapSkip(
        apiKey=os.environ.get("CAPSKIP_API_KEY", "capskip"),
        host=os.environ.get("CAPSKIP_HOST", "127.0.0.1"),
        port=8080,
    )

    async def parse(self, response):
        widget = response.css("altcha-widget")
        source = (widget.attrib.get("challenge")
                  or widget.attrib.get("challengejson")
                  or widget.attrib.get("challengeurl"))
        if not source:
            self.logger.warning("no ALTCHA challenge on %s", response.url)
            return
        field = widget.attrib.get("name", "altcha")

        if source.lstrip().startswith("{"):
            challenge_json = source
        else:
            # Inline request: the page's jar and proxy, no dupefilter.
            meta = {"dont_cache": True, "allow_offsite": True}
            for key in ("cookiejar", "proxy"):
                if key in response.meta:
                    meta[key] = response.meta[key]
            reply = await self.crawler.engine.download_async(
                scrapy.Request(response.urljoin(source), meta=meta,
                               headers={"Referer": response.url})
            )
            if reply.status != 200:
                self.logger.error("challenge endpoint answered %s", reply.status)
                return
            challenge_json = reply.text

        try:
            result = await self.solver.altcha(
                url=response.url, challenge_json=challenge_json)
        except CapSkipError as exc:
            # Expired challenge, unsupported algorithm, or CapSkip unreachable.
            self.logger.error("ALTCHA not solved on %s: %r", response.url, exc)
            return

        form = response.xpath("//form[.//altcha-widget]")
        data = {"email": "[email protected]", field: result["token"]}
        yield form2request(form, data).to_scrapy(
            callback=self.after_submit,
            priority=10,
            meta={"handle_httpstatus_all": True},
        )

    def after_submit(self, response):
        yield {"url": response.url, "status": response.status}

Run it with scrapy runspider and the file name, or drop the class into a project. The spider handles both widget generations, fetches the challenge inside the page’s session, and yields one item per submitted form, with the status the site answered, rejections included. Because the challenge fetch bypasses the scheduler, pages that share one challenge endpoint each get a challenge of their own. The raw endpoint behind the altcha method is documented in the API reference, and the plain Python call without Scrapy around it is covered in the Python ALTCHA guide.

On Scrapy older than 2.14

Before 2.14 there is no download_async, but from Scrapy 2.6 the same inline fetch works through engine.download, awaited with a helper that turns its result into something a coroutine can wait on.

from scrapy.utils.defer import maybe_deferred_to_future

# Scrapy 2.6 to 2.13: the same inline fetch, through engine.download.
reply = await maybe_deferred_to_future(self.crawler.engine.download(
    scrapy.Request(response.urljoin(source),
                   meta={"dont_cache": True, "allow_offsite": True})
))

On those versions FormRequest.from_response builds the submit without a warning, so you can use it in place of form2request: pass the same XPath as formxpath and put the token in formdata. If you route the challenge through the scheduler instead, as an ordinary request with a callback of its own, set dont_filter on it, or the duplicate filter drops every fetch after the first.

On a Scrapy older than 2.13, AsyncCapSkip also needs TWISTED_REACTOR set to the asyncio reactor in settings.py. Projects generated since Scrapy 2.7 already have that line.

Common errors and what they mean

What you seeCauseFix
The first form submits and later ones never do, with no errorThe challenge was yielded as a normal request and the duplicate filter dropped the repeatsFetch it with download_async, or set dont_filter on the request
An ApiException within a moment, but only after the first runThe HTTP cache replayed an old, expired challengeAdd dont_cache to the challenge request
The site says verification failed although the solve succeededThe token waited too long, was reused, or went in the wrong fieldSubmit straight away with a raised priority, once per token, under the widget’s name attribute
A 403 from the challenge endpointThe endpoint wants the session that loaded the formFetch it through Scrapy, as above, not through challenge_url
An ApiException in under a second, every timeThe site uses Argon2id or scrypt, which are refused rather than attemptedNothing to retry; that site needs another route
The crawl stalls while each form is solvedThe synchronous CapSkip class is running on the event loopSwitch to AsyncCapSkip and an async def callback
NoEventLoopError, not currently running on any asynchronous event loop (AsyncLibraryNotFoundError on older installs)The project pins a reactor other than the asyncio oneRemove that TWISTED_REACTOR line, or set it to the asyncio reactor
AttributeError: ‘ExecutionEngine’ object has no attribute ‘download_async’Scrapy is older than 2.14Upgrade, or use engine.download as shown above
IgnoreRequest, filtered offsite request, on the challenge fetchThe challenge endpoint is on another domain and the spider sets allowed_domainsAdd allow_offsite to the challenge request’s meta
A ScrapyDeprecationWarning about from_response on every submitScrapy 2.16 and later deprecate FormRequest.from_responseBuild the submit with form2request
A NetworkException on the first solveCapSkip is not running, or the host and port are wrongStart CapSkip, then check whether it should be in Local mode or Server mode

FAQ

Do I need scrapy-playwright or a headless browser for ALTCHA?

No, you can solve ALTCHA in Scrapy with plain requests. The widget’s whole job is to fetch a challenge, burn some CPU finding a number, and write the result into a field. Scrapy can fetch the challenge, CapSkip finds the number, and form2request writes the field, so no JavaScript has to run anywhere. That keeps the spider as fast and as cheap to run as any other Scrapy crawl.

Can a spider on Scrapyd or a hosted Scrapy platform reach the solver?

Yes. Switch CapSkip to Server mode under connection settings so it listens on a network address instead of loopback, set CAPSKIP_HOST in the spider’s environment, and pass it to the client as the full example does. A hosted platform connects over the same HTTP API as a local spider. Use a static public IP with a firewall rule when the route crosses the internet. The solver stays on hardware you own, so the solve count never changes what you pay.

Why is this in a callback and not a downloader middleware?

Because ALTCHA sits on a form you choose to submit, not on a block page that interrupts the crawl. A middleware is the right home for a challenge that can appear on any response, which is how the reCAPTCHA middleware guide handles it. ALTCHA is part of one step of the flow, and the callback that builds that form already has everything the solve needs.

Can I solve many forms in parallel?

Yes, and Scrapy already does it for you. Every async callback that awaits the solver yields control, so every page Scrapy has downloaded can have its solve in flight at the same time. CONCURRENT_REQUESTS does not cap that, because it limits downloads, not callbacks. On the CapSkip side, Max. Threads in the ALTCHA settings sets how many challenges hash at once, and since that is CPU work, more threads than cores buys nothing. Because each callback fetches its own challenge just before solving, none of them sits in a queue going stale. Do not prefetch a batch of challenges to solve later, for the same reason.

The short version

To solve ALTCHA in Scrapy, read the challenge, challengejson or challengeurl attribute and the name attribute off the altcha-widget element. Fetch the challenge with download_async and dont_cache, so it rides the spider’s session and never meets the duplicate filter. Await AsyncCapSkip’s altcha method with the JSON as it arrived, and yield a form2request submit with the token in its data and a raised priority. One challenge, one token, one submit, and Server mode when the spider runs somewhere else.

One last thing about volume. A crawler that submits thousands of forms solves thousands of challenges, and with a captcha solver running on your own machine that costs CPU time rather than a per-solve fee.