Back to Blog

Containers as a Service for Web Scraping with Scrapeless

Daniel Kim
Daniel Kim

Lead Scraping Automation Engineer

28-Sep-2026

TL;DR:

  • Containers as a service manages the runtime for scraping workers. Your application still defines the target, validates the response, and decides where records belong.
  • Scrapeless Web Unlocker handles the web-access request outside the worker container. The container needs an HTTP client rather than a locally installed browser for this workflow.
  • A completed process is not proof of useful data. Check the target content before writing an accepted record.
  • Runtime secrets and durable storage belong outside the image. A replacement container should start from configuration, not recover credentials from an old filesystem.
  • Free to start. Create a Scrapeless account and use the available free credit to evaluate a small workload.

Introduction: A Container Should Do One Predictable Job

A scraping worker spends much of its lifetime waiting for another system. It sends a request, waits for a page, checks the response, and writes a result. Packaging those steps in a container makes the execution environment repeatable. It does not determine whether the returned page contains the information the application needs.

Containers as a service, or CaaS, gives a team managed infrastructure for deploying and running those containers. The useful design question is where to put each responsibility. The runtime schedules the worker. The worker owns the task. A web-access service handles the target request. Storage retains the evidence after the worker exits.

This tutorial develops a small Python worker around Scrapeless Web Unlocker and shows how to package it for a container platform. The same separation is useful in a competitive pricing pipeline, where a page retrieval is only one stage before data normalization and analysis.

What Containers as a Service Actually Manages

Containers as a service manages container execution while your application retains responsibility for its workload and data. Depending on the platform, the managed layer can cover scheduling, resource allocation, networking, health checks, and scaling.

A Dockerfile describes how to build an image. An image is the packaged application. A container is a running instance. An orchestrator decides where and when instances run. These concepts work together, but a Dockerfile alone does not provision a managed platform. The Dockerfile build model is the starting point for packaging the worker below.

Responsibility Owner in this design Evidence to retain
Schedule a task Your application or job scheduler Task identifier and approved URL
Run the worker Container platform Exit status and resource usage
Request the page Scrapeless Web Unlocker Raw response and service outcome
Validate content Your worker Expected-content check
Save accepted data Your storage integration Record key and write confirmation

CaaS is useful when the same worker must run repeatedly across environments or when task volume requires multiple workers. A single local script may be enough for an occasional export. Choose a managed runtime when its operational benefits justify deployment and monitoring work.

Pipeline at a Glance

The pipeline turns an approved URL into a stored page response with an explicit validation decision. Start with one task per process, then add a queue after the task contract is stable.

The sequence is: task input → Web Unlocker request → response capture → expected-content check → accepted record. The demonstration writes to a mounted output directory. A production implementation should replace that directory with durable storage or a persistent volume suited to the chosen platform.

Scrapeless Web Unlocker sits in the access step. It is not a container scheduler, queue, or database. Keeping that boundary clear makes it possible to change the deployment platform without rewriting the extraction policy.

Prerequisites

You need Python, the Requests package, an active Scrapeless API key, and a public target you are authorized to collect. Set SCRAPELESS_API_KEY, TARGET_URL, and EXPECTED_TEXT in the runtime environment. The last value is a phrase that should occur in the intended page, not a universal challenge detector.

Container execution additionally requires Docker or a compatible build runtime. Deployment requires a container registry and a configured CaaS account or cluster. Authenticated Scrapeless execution and the container build are environment-dependent prerequisites for the example; no successful service response or managed deployment is asserted here.

For the local Python dependency, run python -m pip install requests. The Web Unlocker request contract defines the actor and request envelope used below.

Stage 1: Write a Bounded Page Worker

The worker submits one request and retains its raw response before deciding whether the page is acceptable. Save the following as worker.py.

Note: This block requires your Scrapeless API key and an approved target. The authenticated request has not been executed for this example; validate the response against your account before deployment.

python Copy
import hashlib
import json
import os
import time
from pathlib import Path

import requests

url = os.environ["TARGET_URL"]
expected = os.environ["EXPECTED_TEXT"]
output = Path(os.environ.get("OUTPUT_DIR", "/output"))
output.mkdir(parents=True, exist_ok=True)
response = requests.post(
    "https://api.scrapeless.com/api/v2/unlocker/request",
    headers={"x-api-token": os.environ["SCRAPELESS_API_KEY"]},
    json={
        "actor": "unlocker.webunlocker",
        "input": {"url": url, "method": "GET", "redirect": False},
        "proxy": {"country": "ANY"},
    },
    timeout=120,
)
response.raise_for_status()
body = response.content
capture_id = hashlib.sha256(body).hexdigest()
(output / f"{capture_id}.response").write_bytes(body)
if expected.casefold() not in response.text.casefold():
    raise RuntimeError("Expected page content was not found")
record = {
    "requested_url": url,
    "collected_at_unix": int(time.time()),
    "response_sha256": capture_id,
    "response_bytes": len(body),
    "http_status": response.status_code,
    "validation": "expected_text_present",
}
(output / f"{capture_id}.json").write_text(json.dumps(record, indent=2))
print(json.dumps(record))

The record is an application-defined output, not a claim about Scrapeless response fields. The raw response is retained without assuming that every target produces the same content type. The configured timeout bounds the client wait; it does not define a service-level guarantee.

An expected phrase is a minimal check. For structured data, replace it with a parser that requires the necessary fields and validates their types. A title appearing in an error message should not qualify as a valid product record. The HTTP response semantics describe protocol outcomes; business validation belongs to your application.

Start Scraping with Scrapeless

Power up your web scraping and automation workflow with Scrapeless!
Sign up today and get $5 in free credit — no credit card required.

Claim your free credit now in the Scrapeless Dashboard.

Stage 2: Package the Worker Without Credentials

The image should contain code and dependencies while the runtime supplies secrets. Put this Dockerfile beside worker.py.

Note: Building this configuration requires an installed container runtime and registry access. The image build and container execution remain deployment prerequisites; the configuration is not a captured deployment result.

dockerfile Copy
FROM python:3.12-slim
WORKDIR /app
RUN pip install --no-cache-dir requests
COPY worker.py /app/worker.py
CMD ["python", "/app/worker.py"]

This minimal image deliberately keeps dependency management visible. For a release, resolve and lock the dependency versions in your build environment, scan the resulting image, and deploy the image digest you approved. Do not place an API key in a Dockerfile, build argument, or copied environment file. A runtime secret configuration keeps credential delivery separate from the image.

For a local container check, prepare the environment variables in your shell and create a writable output directory before running the commands below. Passing an environment variable by name avoids writing its value into the command.

Note: These commands require Docker, the files above, and the authenticated runtime configuration. They have not been run against a local Docker engine in this example.

bash Copy
docker build -t scrapeless-page-worker .
mkdir -p output
docker run --rm \
  -e SCRAPELESS_API_KEY -e TARGET_URL -e EXPECTED_TEXT \
  -v "$PWD/output:/output" \
  scrapeless-page-worker

A zero exit code means this worker reached its final print statement. Confirm the raw capture and metadata file both exist, inspect the page content, and check that the configured target matches the intended source before treating the example as accepted.

Stage 3: Deploy as a Job, Then Introduce a Queue

A job is the natural deployment shape for a worker that processes a bounded input and exits. The Kubernetes Job workload represents this execution pattern on Kubernetes-based platforms; other managed container systems expose similar task or job concepts with different configuration.

Choose one platform and configure its image reference, command, secret injection, output destination, and task deadline. Run the same small target set used locally. Compare accepted records, rejected responses, and total service usage before increasing the worker count.

A queue becomes useful when jobs need shared scheduling. Define a task identifier that remains stable when the same scheduled observation is delivered more than once. Include the observation period in that identifier if the purpose is to collect snapshots over time. Otherwise, a URL-only key can incorrectly collapse distinct observations.

Write the accepted record before acknowledging completion. Use the task identifier to prevent duplicate writes. Container replacement must not turn an uncertain storage operation into a second business record.

Stage 4: Measure Accepted Records and Resource Cost

Worker monitoring should distinguish process health from data quality. Track accepted records, rejected responses, queue age, and time spent waiting for the service. CPU usage alone may be a poor scaling signal for an application that mostly waits for network responses.

Keep the initial concurrency small and bound work per target. An internal starting limit of at most three workers per host is a conservative example setting, not a universal website allowance. Site policy and account limits can require a lower value.

Compare container runtime cost with actual Scrapeless usage on the pricing page. Do not infer API consumption from worker uptime: a running container can be idle, while a short job can perform several billable operations.

Protect raw responses according to their contents. Request headers, cookies, and any authorized session material need stricter handling than public page text. Only retain the evidence necessary for the extraction task.

Conclusion: Deploy the Contract You Can Validate

A useful CaaS scraping workflow has a small task contract, a worker that checks its output, and storage that survives process termination. Start with the single-page worker, verify its response handling with your account, and run the same code inside a container before adding orchestration.

The next deployment milestone is an accepted record with its source and capture evidence. Worker count comes after that result is reproducible.

Ready to Build Your Web Data Pipeline?

Join developers working on web data collection in Discord and Telegram.

Create a Scrapeless account and adapt the workflow to your own approved data sources.

FAQ

Q: Is scraping from a container legally different from running a local script?

Container deployment does not change the access permissions or data-use obligations for the target. Review applicable terms and rules for the sources and data involved.

Q: Does a CaaS platform provide the proxy for this worker?

This worker delegates its target request to Web Unlocker and uses the documented proxy configuration in that request. A container's own outbound IP is not automatically a residential proxy.

Q: What should happen when the response is an access-denied page?

Reject the page as task data and retain a minimal diagnostic record. Check the target scope and supported access path before authorizing another collection job.

Q: What changes when the website's HTML changes?

Update and validate the extraction contract. Packaging code in a container does not make selectors or expected content immune to page changes.

Q: How many workers should run at once?

Begin with a small bounded workload and measure accepted output per target. The example's ceiling of three workers per host is an application setting, subject to stricter target and account limits.

Q: Does this workflow require an AI agent?

No. The Python worker and container commands run without a language model. An agent may create tasks, but it should not replace the worker's deterministic content checks.

At Scrapeless, we only access publicly available data while strictly complying with applicable laws, regulations, and website privacy policies. The content in this blog is for demonstration purposes only and does not involve any illegal or infringing activities. We make no guarantees and disclaim all liability for the use of information from this blog or third-party links. Before engaging in any scraping activities, consult your legal advisor and review the target website's terms of service or obtain the necessary permissions.

Most Popular Articles

Catalogue