Web Scraping With PHP: cURL, DOMXPath, and Rendering

Web Scraping With PHP

Scrapeless Universal Scraping API returns fetched or rendered page content to PHP applications through an authenticated HTTP call.

TL;DR

  • PHP already has a capable static-page stack. The cURL extension fetches responses while DOMDocument and DOMXPath turn HTML into structured queries.
  • Extension availability must be checked first. CLI installations can omit curl or DOM even when the PHP runtime itself is present.
  • HTML parser behavior depends on the runtime. Modern HTML and malformed markup can produce warnings or a tree that differs from a browser DOM.
  • XPath works best when scoped to a record node. Relative queries prevent the first page-wide title or price from leaking into every row.
  • Rendered acquisition handles the JavaScript boundary. The PHP parser can stay unchanged when another layer returns the finished markup.

What PHP Needs for a Scraper

Web scraping with PHP can be built from extensions that many web servers already carry: cURL for HTTP, DOM for parsing, DOMXPath for selection, and JSON for output. Confirm those modules in the exact CLI or container runtime that will execute the job. A web-hosting image and its command-line image may not expose the same extensions.

The PHP cURL manual defines the session handle used to configure a request. Set the target URL, return-transfer behavior, redirect policy, headers, and timeouts explicitly. Inspect both the returned body and cURL metadata before handing bytes to a parser.

Keep fetching and parsing in separate functions. That division makes it possible to test XPath against a saved fixture, test transport against a controlled endpoint, and switch to rendered acquisition without rewriting the extraction rules.

Select the Parser for the Runtime

DOMDocument remains common across supported PHP installations, but its loadHTML method parses using older HTML rules. The official DOMDocument documentation recommends the newer HTML document API in runtimes that provide it because modern browser parsing follows HTML5 rules. Check the deployment version before choosing the constructor.

PHP componentJobFailure to surface
cURLFetch page bytesNetwork error, status, redirect, wrong content type
DOMDocumentBuild an HTML treeWarnings, malformed input, parser-version differences
DOMXPathSelect records and fieldsEmpty or page-wide matches
JSON extensionSerialize outputEncoding and invalid value errors
Application schemaNormalize fieldsMissing identity, ambiguous numbers, duplicate keys

Fetch and Parse One Public Page

This block is a local-runtime prerequisite because PHP is not installed in the current workspace. Run it in a PHP environment where curl and DOM are enabled. The example keeps network and parser errors distinct and reads the heading and first link from Example Domain.

<?php
$url = 'https://example.com/';
$handle = curl_init($url);

curl_setopt_array($handle, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_FOLLOWLOCATION => true,
    CURLOPT_CONNECTTIMEOUT => 15,
    CURLOPT_TIMEOUT => 25,
]);

$html = curl_exec($handle);
if ($html === false) {
    throw new RuntimeException(curl_error($handle));
}

$status = curl_getinfo($handle, CURLINFO_RESPONSE_CODE);
$finalUrl = curl_getinfo($handle, CURLINFO_EFFECTIVE_URL);
curl_close($handle);

if ($status < 200 || $status >= 300) {
    throw new RuntimeException('Unexpected HTTP status');
}

libxml_use_internal_errors(true);
$document = new DOMDocument();
$document->loadHTML($html);
libxml_clear_errors();

$xpath = new DOMXPath($document);
$heading = $xpath->query('//h1')->item(0);
$anchor = $xpath->query('//a[@href]')->item(0);

if ($heading === null || $anchor === null) {
    throw new RuntimeException('Expected page identity is missing');
}

$record = [
    'title' => trim($heading->textContent),
    'link' => $anchor->getAttribute('href'),
    'source' => $finalUrl,
];

echo json_encode($record, JSON_PRETTY_PRINT | JSON_THROW_ON_ERROR);

Internal libxml warnings are collected around parsing rather than hidden globally for the lifetime of the process. In a service, restore the previous error setting after the parse. The response body should also have a size limit suited to the expected page type before it reaches DOMDocument.

Write Relative XPath Queries

DOMXPath can express field relationships that are awkward in a single CSS chain. Select repeated cards with one absolute query, then query fields relative to the current card using a leading dot. Without that dot, a loop can select the first matching title in the entire document for every record.

XPath returns node lists, so every item lookup can be absent. Treat missing identity fields as rejection. Map genuinely optional values to null, and keep normalization outside the XPath expression. Complex string manipulation inside a selector is difficult to test and easy to tie to one presentation.

  • Anchor records to a durable container. Use semantic attributes or stable URL shapes instead of layout depth.
  • Resolve links against the final URL. Redirects can change the base path used by relative references.
  • Check the document identity. A heading or canonical link should confirm that the expected page arrived.
  • Keep raw and normalized fields separate. Locale-sensitive prices and dates need source-aware conversion.

Know Where Plain PHP Stops

cURL downloads the server response and DOMDocument parses it. Neither runs client-side JavaScript. If the response contains a root container but the browser later displays rows inside it, the PHP parser is operating correctly on incomplete source data.

Diagnose the boundary before changing selectors. Search the raw body for a visible value, inspect the network response that supplied it, and decide whether a permitted structured source or a rendered page is appropriate. A render service can return finished HTML while the PHP extraction function stays deterministic.

Paginate Without Losing State

For link pagination, select the next anchor and resolve it against the final page URL. For page numbers, keep a maximum page budget. For cursors, store the continuation token exactly as returned. Maintain a visited set so alternate query order or tracking parameters cannot create a cycle.

An empty item list is not a complete stopping rule. It may mean a changed XPath, a wrong locale, a consent page, or a client-rendered shell. Require an end marker or absent continuation signal, and treat identity or count mismatches as a failed page.

Validate the Response and the Dataset

The HTTP semantics specification defines status behavior, yet a successful status only says that a response was served. Check final URL, content type, body size, identity marker, container count, and required fields before publishing records.

Emit compact batch metrics instead of page bodies: pages requested, pages accepted, XPath misses, records emitted, records rejected, and durations. Keep credentials out of exception messages. Store test fixtures in controlled project data rather than copying arbitrary production responses into logs.

Control Memory and Output Boundaries

DOMDocument builds a tree in memory, so response limits matter before parsing begins. Set a maximum body size that matches the expected page class and reject unexpected downloads rather than letting one document consume the worker. For a long crawl, parse one page, convert it to small records, release the document, and then move to the next page.

Write records incrementally as JSON Lines, database batches, or messages instead of collecting the full crawl in one PHP array. Incremental output ties memory use to the current page and makes partial progress observable. The writer should accept only validated records; parser warnings, page diagnostics, and rejected rows belong in a separate operations channel.

Keep acquisition metadata at the batch boundary. The final URL, content type, parse method, and acquisition time explain what produced a record without mixing diagnostic fields into the business schema. This also makes a later parser migration easier to audit.

Make that contract visible in a small integration test before any scheduled collection begins.

Conclusion

Web scraping with PHP needs no large framework for static pages: cURL fetches, DOM builds the tree, and DOMXPath selects fields. The durable version checks extensions, surfaces parser warnings, scopes XPath to each record, and validates page identity. If scripts create the data, replace the acquisition layer and keep the parser contract intact.

Ready to Extend a PHP Data Collector?

Connect rendered acquisition to the same cURL boundary and DOMXPath parser, then verify one bounded public-data job before scheduling it.

Sign up today and get $5 in free credit — no credit card required.

Claim Your $5 Credit →

FAQ

Can PHP scrape websites without Composer packages?

Yes. The curl, DOM, DOMXPath, and JSON extensions cover a static-page scraper. Availability depends on the installed runtime, so check extensions before deployment.

Why does DOMDocument show HTML warnings?

Real HTML can be malformed or follow browser rules that differ from the parser used by the runtime. Capture libxml warnings around the parse and test the resulting tree against fixtures.

Can PHP cURL execute JavaScript?

No. cURL retrieves the server response and does not run a browser. Use rendered acquisition when scripts create the required elements.

Should a PHP scraper use XPath or CSS selectors?

DOMXPath is built into the DOM stack and is expressive for relationships and attributes. A CSS-to-XPath library can be useful, but stable field scoping matters more than the selector syntax.

Is web scraping with PHP legal?

PHP does not change the permission analysis. Use authorized public sources, review terms and robots rules, respect technical restrictions, and obtain legal advice for sensitive or regulated collection.

References