Web Scraping With Ruby: Net::HTTP and Nokogiri Guide

Web Scraping With Ruby

Scrapeless Universal Scraping API provides Ruby clients with fetched or rendered page content through an authenticated HTTP request.

TL;DR

  • Ruby separates fetching and parsing with a small toolset. Net::HTTP handles transport and Nokogiri turns HTML into a document queried with CSS or XPath.
  • Nokogiri is a parser, not a browser. It extracts what exists in the supplied markup and does not execute page JavaScript.
  • Row-scoped queries protect data quality. Calling at_css on the current card avoids repeating the first page-wide field across every record.
  • Absolute URLs need an explicit base. Resolve each relative link against the final response URI before storing it.
  • A rendered fetch can preserve the Ruby model. Change how HTML is obtained while keeping Nokogiri selectors and normalization methods.

Build Ruby Scraping as Two Layers

Web scraping with Ruby is easiest to maintain when the HTTP client and HTML parser have separate responsibilities. Net::HTTP sends the request and returns response metadata plus bytes. Nokogiri parses those bytes and exposes CSS and XPath selection. A normalization method maps page strings into application records.

The Ruby Net::HTTP documentation covers HTTPS, proxy configuration, request methods, and response handling. Reuse a configured session for a sequence of requests to the same host, and keep timeouts explicit so a scheduled worker has a known upper bound.

Nokogiri’s official parsing tutorial accepts strings, files, and network sources. For a crawler, passing a response body directly keeps transport behavior visible and lets the parser operate on saved fixtures during tests.

Decide Between Static and Rendered Acquisition

Source behaviorRuby pathReason
Required values exist in response HTMLNet::HTTP plus NokogiriDirect and low overhead
A JSON response supplies the pagePermitted endpoint plus JSONStructured fields avoid DOM interpretation
Scripts create the useful DOMRendered acquisition plus NokogiriParser receives finished markup
Clicks or scrolling define stateBrowser automationThe task is interaction, not only document parsing

Diagnose with source bytes rather than the inspector alone. If a value visible on screen is absent from the response, a different selector cannot reveal it. Decide whether the acquisition layer should render the page or use an authorized structured source.

Fetch and Parse a Page With Ruby

This example is a local-runtime prerequisite because Ruby is not installed in the current workspace. Install Nokogiri in the project, run the block locally, and keep the target bounded to a page you are allowed to collect. The code reads the heading and first link from Example Domain.

require "net/http"
require "nokogiri"
require "json"
require "uri"

source = URI("https://example.com/")

response = Net::HTTP.start(
  source.host,
  source.port,
  use_ssl: true,
  open_timeout: 15,
  read_timeout: 20
) do |http|
  http.get(source.request_uri)
end

unless response.is_a?(Net::HTTPSuccess)
  raise "Unexpected HTTP status: #{response.code}"
end

document = Nokogiri::HTML(response.body)
heading = document.at_css("h1")
anchor = document.at_css("a[href]")

raise "Expected page identity is missing" unless heading && anchor

record = {
  title: heading.text.strip,
  link: URI.join(source, anchor["href"]).to_s
}

puts JSON.pretty_generate(record)

css returns a node set, while at_css returns the first matching node or nil. That distinction is useful in a schema: use at_css for a single field and check absence explicitly; use css for repeated records.

Keep Queries Inside the Current Card

Select every record container from the document, then call at_css or at_xpath on the current node. The Nokogiri search tutorial demonstrates both selector styles. Scoped queries keep one missing card field from falling back to the first match elsewhere on the page.

Prefer selectors tied to meaning: a stable attribute, semantic element, or durable link form. Avoid class chains copied from visual wrappers. Store the canonical source key with each record and reject rows that lack it. Optional fields should remain nil until a normalization rule has enough information to convert them.

  • Trim presentation whitespace at the field boundary. Preserve internal text where spacing carries meaning.
  • Resolve links as soon as they are extracted. Relative paths depend on the final response location.
  • Validate a known page marker. A title or canonical path can distinguish content from an interstitial.
  • Keep extraction free of business defaults. Missing availability should not become false, zero, or “unknown” without a defined rule.

Model Pagination as a Small Enumerator

Ruby can express pagination as an Enumerator that yields one page result at a time. The iterator carries the next URL or cursor, checks that it stays inside the allowed scope, and stops when the source removes its continuation signal. This keeps discovery separate from storage.

Maintain a set of canonical continuation identities to prevent loops. Add a page budget for bounded jobs. Do not interpret an empty node set as proof of completion unless the source contract also says there is no next page.

Separate Normalization From HTML

Nokogiri methods should return source-shaped values. A separate object can parse numbers, currencies, dates, and availability labels. This division makes locale rules testable and preserves raw evidence when a conversion fails.

Use hashes for a small script and a class or immutable value object when the schema grows. Required identifiers should raise or reject at construction. Optional fields can remain nil. Add provenance such as canonical URL and acquisition time beside the record so downstream changes can be traced to their source.

Validate More Than the Status

The HTTP semantics specification defines successful status classes, but a successful response can still be the wrong page. Check final location, content type, body size, page marker, and plausible record count before accepting a batch.

Log compact facts: requested host, response class, parse duration, containers matched, records accepted, and rejection category. Do not place full production pages or credentials into ordinary logs. Use controlled fixtures for selector tests.

Manage Cookies and Sessions Deliberately

A sequence of public pages may still use cookies for locale, consent state, or navigation continuity. Net::HTTP does not provide a complete browser cookie jar by itself, so decide whether the workflow needs session state before adding it. For a stateless catalog, independent requests may be clearer. For a permitted stateful flow, use a maintained cookie-jar abstraction and keep one session confined to one logical task.

Do not share authenticated or personalized cookies across unrelated jobs. Session state changes page identity and can expose data outside the intended public scope. Record the locale and state class used for each batch, but keep cookie values and credentials out of logs and output records.

Test Selectors With Purpose-Built Fixtures

Save a small representative HTML fixture for a normal page, one with an optional field missing, and one that should fail the page-identity check. Parser tests can then assert exact record values without depending on a live site. A fixture should be reviewed, stripped of unnecessary personal data, and stored with the code that owns the selector.

Test the acquisition method separately against a bounded endpoint. Confirm status, final location, content type, and the expected marker. An integration test can then join acquisition and parsing for one page and validate the serialized schema. This layered test design makes a failure local: transport, page identity, selector, normalization, or output.

Attach provenance to accepted batches, including canonical source URL, acquisition time, and parser revision. These fields let a downstream analyst distinguish a real source change from a code deployment or a regional page variant.

Conclusion

Web scraping with Ruby stays compact when Net::HTTP fetches, Nokogiri parses, and a separate model normalizes. Scope queries to each record, resolve links against the response URI, and reject unexpected page identities. For a client-rendered target, preserve the Ruby extraction layer and change only how the finished HTML reaches it.

Ready to Add Rendered HTML to a Ruby Collector?

Connect Scrapeless to the fetch boundary, retain Nokogiri and your Ruby value objects, and validate a small authorized workflow.

Sign up today and get $5 in free credit — no credit card required.

Claim Your $5 Credit →

FAQ

Is Nokogiri a web scraper?

Nokogiri is an HTML and XML parser. Net::HTTP or another acquisition layer fetches the content, and Nokogiri then selects nodes from that content.

Can Nokogiri execute JavaScript?

No. Nokogiri parses supplied markup and has no browser event loop. Use rendered acquisition when page scripts create the required elements.

Should Ruby scrapers use CSS or XPath?

Nokogiri supports both. CSS is concise for common selectors, while XPath is useful for relationship and text queries; durability and row scoping matter more than the syntax.

How should Ruby represent a missing scraped value?

Use nil for a genuinely optional field and reject the record when a required identity is absent. Do not convert absence into a plausible business value during extraction.

Is web scraping with Ruby legal?

Ruby does not grant permission to collect data. Use authorized public sources, review terms and robots instructions, respect access controls, and seek legal advice for sensitive or high-impact uses.

References