Back to Blog

Web Scraping with Java: Jsoup, Playwright, and Cloud Browser Workflows

Ava Wilson
Ava Wilson

Expert in Web Scraping Technologies

10-Oct-2026

TL;DR:

  • Jsoup parses the HTML supplied to it. It does not run a page's JavaScript to create missing records.
  • Playwright Java supplies a browser acquisition path. The rendered HTML can feed the same Jsoup extraction contract used for a static page.
  • Agent Browser moves browser execution to the cloud. Java still owns selectors, record validation, pagination scope, and export.
  • A CSV file is useful only when its records identify the intended source. Preserve stable record IDs and resolved URLs, and reject missing required fields.

Java web scraping starts with a practical question: does the server's HTML already contain the records, or does a browser need to create them? Choosing the acquisition path correctly avoids writing a selector for a document that never contained the data.

This workflow uses Jsoup for extraction and Playwright Java for rendered-page acquisition. Both paths produce the same record shape. A dedicated Scrapeless Agent Browser connection is the cloud option when the browser workload should run outside your application host.

The examples target a controlled article listing with stable attributes. Replace that contract only after inspecting your permitted source. Java compilation and browser execution require the runtime and credentials listed below; no live Java capture is claimed in this guide.

Choose Jsoup or a Browser from the Page's Actual Content

Choose Jsoup when the required records are present in the retrieved HTML. Choose a browser when the permitted page requires JavaScript or interaction before those records exist.

Path HTML supplied to the parser Suitable starting task Responsibility you keep
Jsoup HTTP fetch Server-returned document Static article or catalogue listing Source inspection and extraction
Local Playwright Java Browser-rendered document Controlled dynamic-page development Browser runtime and resource management
Agent Browser with Playwright Java Cloud-browser document A dedicated remote browser workflow Connection lifecycle, selectors, and output checks

Do not choose the path from the framework used to build the site. A JavaScript application can serve useful HTML, while a simple-looking listing may insert records after load. Inspect the specific page and required fields.

HTTP representation semantics describes the response a client receives. The page state produced later by a browser is a different acquisition observation.

Prerequisites and Dependencies

This project needs a JDK and Maven, plus the Jsoup and Playwright Java dependencies. The shown project targets JDK 17; it pins Jsoup 1.23.2 and Playwright 1.63.0 rather than reusing versions from an older tutorial.

For local rendering, install the browser runtime required by your Playwright version in your own project environment. For remote rendering, obtain a dedicated Agent Browser connection through the current account setup and keep its complete endpoint in SCRAPELESS_CDP_URL.

The target must be public or otherwise authorized. TARGET_URL is a permitted article-listing URL whose structure you have inspected. The example's selector contract is main article[data-id], with a heading and link inside each article.

Note: The Java toolchain, dependency installation, and browser runtime are execution prerequisites. The snippets were checked against current library interfaces but were not compiled or executed in the available verification environment. The remote path also requires a real dedicated Scrapeless connection.

Save this Maven configuration as pom.xml and the class below as src/main/java/ArticleCollector.java:

xml Copy
<project xmlns="http://maven.apache.org/POM/4.0.0">
  <modelVersion>4.0.0</modelVersion>
  <groupId>example</groupId>
  <artifactId>article-collector</artifactId>
  <version>1.0.0</version>
  <properties>
    <maven.compiler.source>17</maven.compiler.source>
    <maven.compiler.target>17</maven.compiler.target>
    <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
  </properties>
  <dependencies>
    <dependency>
      <groupId>org.jsoup</groupId>
      <artifactId>jsoup</artifactId>
      <version>1.23.2</version>
    </dependency>
    <dependency>
      <groupId>com.microsoft.playwright</groupId>
      <artifactId>playwright</artifactId>
      <version>1.63.0</version>
    </dependency>
  </dependencies>
  <build>
    <plugins>
      <plugin>
        <groupId>org.apache.maven.plugins</groupId>
        <artifactId>maven-compiler-plugin</artifactId>
        <version>3.10.1</version>
      </plugin>
    </plugins>
  </build>
</project>

Keep the compiler target and dependency versions in this project so another environment can reproduce the setup before any collection job is enabled.

Configure One Extraction Contract for Both Paths

The extraction contract defines an accepted article record independently of the transport. This example requires a nonempty data-id, an h2 heading, and a link that resolves to an HTTP or HTTPS URL.

These are controlled-example attributes, not claims about a public site's current selectors. For your source, prefer stable record attributes or durable URL patterns when available. Do not reuse a decorative class simply because it matched one capture.

Keep required and optional fields separate. A missing identifier rejects the example record. An optional summary can remain absent without changing the record's identity. Make that policy explicit before exporting.

Basic Implementation: Fetch, Extract, and Export

The class below acquires HTML through the selected path, extracts article records with Jsoup, and writes a UTF-8 CSV file. Its remote branch uses an account-supplied CDP endpoint rather than inventing a Java Scrapeless SDK method.

Note: This complete Java example requires the pinned dependencies, JDK, Maven setup, and a permitted target. The local branch needs a browser runtime; the remote branch needs SCRAPELESS_CDP_URL. No output below is presented as a live execution capture.

java Copy
import com.microsoft.playwright.*;
import com.microsoft.playwright.options.WaitUntilState;
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.net.URI;

public class ArticleCollector {
  private static String csv(String value) {
    return "\"" + value.replace("\"", "\"\"") + "\"";
  }

  public static void main(String[] args) throws Exception {
    if (args.length != 2) throw new IllegalArgumentException("mode and URL required");
    String mode = args[0];
    String target = args[1];
    Document doc;
    if (mode.equals("static")) {
      doc = Jsoup.connect(target).get();
    } else {
      if (!mode.equals("local") && !mode.equals("remote"))
        throw new IllegalArgumentException("Unsupported acquisition mode");
      try (Playwright playwright = Playwright.create()) {
        String endpoint = System.getenv("SCRAPELESS_CDP_URL");
        if (mode.equals("remote") && (endpoint == null || endpoint.isBlank()))
          throw new IllegalArgumentException("Dedicated CDP endpoint required");
        Browser browser = mode.equals("remote")
            ? playwright.chromium().connectOverCDP(endpoint)
            : playwright.chromium().launch();
        try {
          Page page = browser.newPage();
          page.navigate(target, new Page.NavigateOptions()
              .setWaitUntil(WaitUntilState.DOMCONTENTLOADED));
          page.locator("main article[data-id] a[href]").first().waitFor();
          doc = Jsoup.parse(page.content(), page.url());
        } finally {
          browser.close();
        }
      }
    }
    StringBuilder output = new StringBuilder("id,title,url\r\n");
    int accepted = 0;
    for (Element article : doc.select("main article[data-id]")) {
      Element heading = article.selectFirst("h2");
      Element link = article.selectFirst("a[href]");
      String id = article.attr("data-id").strip();
      if (id.isEmpty() || heading == null || link == null) continue;
      String title = heading.text().strip();
      String url = link.attr("abs:href");
      URI resolved = URI.create(url);
      if (title.isEmpty() || resolved.getHost() == null ||
          !("https".equals(resolved.getScheme()) || "http".equals(resolved.getScheme())))
        continue;
      output.append(csv(id)).append(',').append(csv(title)).append(',')
          .append(csv(url)).append("\r\n");
      accepted++;
    }
    if (accepted == 0) throw new IllegalStateException("Content contract not satisfied");
    Files.writeString(Path.of("articles.csv"), output, StandardCharsets.UTF_8);
    System.out.println("Accepted article records: " + accepted);
  }
}

Compile the project and copy its dependencies with Maven, then run ArticleCollector using a classpath containing target/classes and the dependency directory. Select static, local, or remote and supply your inspected target URL.

Note: These POSIX-shell commands require the Java toolchain and dependencies above. Set TARGET_URL to the permitted source. Compilation and collection remain execution prerequisites for this example.

bash Copy
mvn -q dependency:copy-dependencies compile
java -cp 'target/classes:target/dependency/*' ArticleCollector static "$TARGET_URL"

Use a semicolon as the classpath separator in a Windows environment. Choose the local or remote mode only after completing its browser-runtime or dedicated-session setup.

The code uses abs:href so a relative link is interpreted against the captured document's base URL. relative URI resolution explains why a path by itself is not a complete source identity. Retain the acquired page URL separately from each discovered article URL in your application's capture manifest.

Render a Dynamic Page with Playwright Java

The browser path waits for a task-specific element before collecting page HTML. domcontentloaded is the navigation milestone; the article locator establishes the separate condition that the example's records are present.

Do not treat either condition as proof that every record has arrived. A listing may load more rows only after an explicit action. Inspect whether the first batch is complete for the requested task before accepting it.

The Playwright browser connection methods distinguishes its native protocol from CDP. A CDP connection is for Chromium-based browsers and has different capabilities from a native Playwright protocol connection.

Keep timeouts and source readiness conditions in project configuration when developing a real collector. A source-specific visible marker is more useful than waiting for unrelated analytics traffic to become idle.

Start Scraping with Scrapeless

Power up your web scraping and automation workflow with Scrapeless!
Sign up today and get $5 in free credit — no credit card required.

Claim your free credit now in the Scrapeless Dashboard.

Where Local Java Scrapers Stop

A local Java parser cannot supply missing rendered content, and a local browser does not remove the need to manage browser resources and acquisition state. Those limits determine when the workflow needs a cloud browser.

Scrapeless Agent Browser supplies the remote browser layer for this Java workflow. Obtain a dedicated connection through the Agent Browser Playwright setup, then pass the complete connection endpoint to the remote branch.

The connection URL may contain credentials. Keep it in the runtime environment, do not log it, and close only the dedicated session allocated for this job. Connecting to a shared browser and closing it would affect other work using that session.

The cloud path changes where acquisition runs. It keeps the same HTML parsing and record validation in Java. A challenge document or incorrect page state should still fail the content contract rather than produce a successful export.

Pagination should follow the source's inspected navigation contract and a bounded collection scope. Discover a next-page URL or permitted interaction from the actual source instead of guessing a page-number parameter.

Preserve the current page identity, discovered next link, and visited-page set. Confirm the link belongs to the approved source scope before navigating. Stop when the task's page budget is reached, a repeated URL appears, or the source establishes that no further page exists.

An absent next link can mean the listing ended or the source failed to render its navigation. Use the page's content and empty-state evidence to distinguish those outcomes. The single-page class above deliberately leaves that decision to the application rather than implying a universal pagination loop.

Export and Validate the Records

The export step should preserve accepted record identity and encode the chosen format correctly. The CSV field quoting rules accounts for delimiter and quote handling; the example doubles embedded quotes and quotes each field.

CSV syntax is separate from content validation. Reopen the output with the intended downstream parser and check required fields, unique identifiers, and source links. Use a spreadsheet import policy that keeps untrusted text as data when a workbook is the consumer.

The CSV header describes the contract: id, title, and url. This is a normative example schema, not a captured dataset. A source with no matching records causes the example to stop; it does not silently label that result as a valid empty listing.

Troubleshoot the Acquisition Boundary

Missing records should be diagnosed against the acquired document and extraction contract. An empty CSV alone does not identify the cause.

Symptom Inspect Likely engineering action
Static HTML lacks the required records Raw response and source readiness Select the permitted browser path
Browser renders unrelated content Final page and page state Correct the starting URL or interaction
Some records lack IDs or titles Source markup and required fields Update or narrow the contract
Relative links export incorrectly Captured base URL and link attributes Resolve links against the correct document
Remote connection cannot be established Endpoint type and dedicated session Correct account and connection configuration

The existing Jsoup-focused Java workflow covers the parser-oriented acquisition approach. This article adds a shared static, local-browser, and cloud-browser contract without replacing that published page.

Conclusion

Java web scraping works best when acquisition and extraction are separate decisions. Use Jsoup for the document you actually have, obtain rendered HTML when the source requires it, and validate the same record contract before export.

Complete the toolchain and dedicated-session prerequisites, test a permitted single page, and then add source-specific pagination with a bounded scope.

Ready to Connect a Java Collector to a Cloud Browser?

Build the browser path with Scrapeless and evaluate its resource needs against the current pricing. Discuss Java workflow questions on Telegram.

FAQ

Q: Can Jsoup execute a website's JavaScript?

Jsoup parses supplied HTML and does not execute page JavaScript. Use a browser acquisition path when the required records exist only after rendering.

Q: Is Java web scraping permitted on any public page?

Public visibility does not establish permission for every collection or use. Review source terms, robots exclusion rules, applicable requirements, and the project's authorization before collecting.

Q: Does this Java workflow require a separate proxy?

A static or local browser client needs whatever permitted routing its source requires. The cloud path uses the allocated Agent Browser configuration; configure its egress through the current account setup.

Q: What should the collector do with a challenge or access-denied document?

The collector should reject unsuitable content and preserve the acquisition reason. A cloud browser connection does not replace the source-content check.

Q: Why can the same selector return no rows after a page update?

The source may have changed markup, rendering behavior, or page identity. Reinspect those observations and the final URL before changing the selector or accepting an empty result.

Q: How much concurrency should a Java pilot use?

Use a bounded pilot with a conservative per-host cap, such as three workers for this example's policy. The source's collection rules and observed resource use govern expansion.

Q: Can this collector run without an AI agent or alongside Selenium?

The Java workflow can run as deterministic code without an AI agent. An existing Selenium project can keep its browser acquisition layer and feed captured HTML into the same validation and export contract.

At Scrapeless, we only access publicly available data while strictly complying with applicable laws, regulations, and website privacy policies. The content in this blog is for demonstration purposes only and does not involve any illegal or infringing activities. We make no guarantees and disclaim all liability for the use of information from this blog or third-party links. Before engaging in any scraping activities, consult your legal advisor and review the target website's terms of service or obtain the necessary permissions.

Most Popular Articles

Catalogue