Web Scraping With Go: net/http and goquery Workflow

Web Scraping With Go

Scrapeless Universal Scraping API supplies Go services with fetched or rendered page content through a standard authenticated HTTP call.

TL;DR

  • Go separates transport from parsing cleanly. The standard net/http package handles requests while goquery turns response HTML into a CSS-queryable document.
  • One shared client should serve a bounded workload. Connection reuse belongs in a long-lived http.Client, while per-request contexts set clear cancellation boundaries.
  • Goroutines need host-aware limits. A semaphore or worker pool protects both the target and the scraper from uncontrolled fan-out.
  • Parsing success is not page success. Validate status, final URL, content type, and a page-identity marker before accepting records.
  • Rendered HTML can preserve the Go parser. When client-side JavaScript creates the data, change the acquisition layer and keep goquery extraction unchanged.

Why Go Fits Long-Running Scrapers

Web scraping with Go suits services that fetch many independent pages and pass structured records into an existing backend. The language has an HTTP client in its standard library, explicit error returns, lightweight concurrency, and straightforward deployment as a single compiled program. Those traits matter most after the first successful selector.

A useful Go design has three packages or layers: acquisition returns bytes plus response metadata, parsing converts bytes into page-shaped values, and normalization produces domain records. Each layer can be tested without the others. A saved HTML fixture tests selectors; an httptest server tests transport behavior; schema tests check the final output.

The Go net/http documentation notes that callers must close response bodies. Closing is not housekeeping trivia: it allows the transport to reuse connections and keeps a crawler from leaking resources across a long queue.

Assemble the Fetch-and-Parse Stack

The standard client handles HTTP, and goquery provides a jQuery-like selection API over Go’s HTML parser. The goquery project documentation shows the same core pattern used in production code: request a page, pass the response body to NewDocumentFromReader, then scope field queries to each repeated element.

LayerGo componentResponsibility
Transporthttp.ClientTimeouts, redirects, connections, headers
Cancellationcontext.ContextDeadline and request lifetime
ParsinggoqueryHTML tree and CSS selectors
CoordinationWorker pool or semaphoreBounded parallel fetches
Outputencoding/jsonStable records and line-oriented export

Do not create a new client for every URL. A shared client reuses its transport and connection pool. Give the client a finite timeout and give each request a context. This makes cancellation visible in the function signature and prevents a worker from occupying a queue slot forever.

Write a Focused Go Extractor

The following example is a local-runtime prerequisite because this workspace does not include a Go toolchain. Its API shape follows the official net/http and goquery documentation. Run it in a module where goquery is installed, and point it only at pages you are permitted to collect.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "net/http"
    "time"

    "github.com/PuerkitoBio/goquery"
)

type Record struct {
    Title string `json:"title"`
    URL   string `json:"url"`
}

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
    defer cancel()

    req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://example.com/", nil)
    if err != nil { panic(err) }

    client := &http.Client{Timeout: 25 * time.Second}
    res, err := client.Do(req)
    if err != nil { panic(err) }
    defer res.Body.Close()

    if res.StatusCode < 200 || res.StatusCode >= 300 {
        panic(fmt.Sprintf("unexpected HTTP status: %s", res.Status))
    }

    doc, err := goquery.NewDocumentFromReader(res.Body)
    if err != nil { panic(err) }

    href, _ := doc.Find("a").First().Attr("href")
    out, _ := json.MarshalIndent(Record{
        Title: doc.Find("h1").First().Text(),
        URL: href,
    }, "", "  ")
    fmt.Println(string(out))
}

A production extractor should return errors instead of calling panic; the short program keeps the lifecycle visible. The important sequence is unchanged: create a contextual request, perform it with a shared client, close the body, reject unexpected status, parse, then extract.

Keep Selectors and URLs Honest

goquery selectors should start at the record container. Calling doc.Find repeatedly inside a loop can accidentally select the first page-wide match for every row. Use the callback’s current selection and search downward. This makes a missing title local to one record rather than a silent duplication across the dataset.

Relative URL resolution belongs in the acquisition or normalization layer. The standard net/url package resolves a reference against the final page URL; string concatenation does not understand path traversal, query replacement, or protocol-relative links. The WHATWG URL Standard documents the behavior browsers follow.

  • Choose stable attributes. Prefer semantic identifiers, durable data attributes, and link patterns over generated layout classes.
  • Record match counts. A sudden zero or unexpected surge should fail validation before storage.
  • Preserve absence. Use pointers or explicit validity fields when an optional value differs from an empty value.
  • Separate raw and normalized values. Keep the source text when currency, locale, or unit parsing can be ambiguous.

Bound Concurrency With a Worker Budget

Goroutines make concurrency easy to start and easy to overuse. Production web scraping with Go should cap work per host, carry cancellation through every request, and close the result channel only after all workers finish. A small worker pool is easier to observe than one goroutine per discovered URL.

Bounded work also protects local resources. Response bodies, parsed trees, JSON buffers, and queued URLs all consume memory. Streaming results to storage or a downstream channel keeps the process size tied to the worker count rather than to the entire crawl.

Concurrency does not replace source pacing or authorization. Use a schedule and volume the target permits. The Robots Exclusion Protocol standard defines a machine-readable way for site operators to express crawler preferences, while terms and applicable law can impose additional obligations.

Model Pagination as State

A Go crawler should represent continuation explicitly. A page-number loop stores the next number. A link-driven crawler resolves the next anchor. A cursor-driven source stores the exact cursor returned with the response. In each case, the state belongs in a typed value passed to the next request, not reconstructed from incidental output.

Keep a set of visited canonical URLs or cursors to prevent cycles. Stop on an explicit end condition and set a maximum page budget for bounded jobs. If the source returns no records but still exposes a continuation token, treat that mismatch as a validation error rather than guessing that the crawl is complete.

Validate Responses Before Publishing Records

An HTTP status is one signal, not proof of page identity. The HTTP semantics specification defines the meaning of success and error status classes, but a successful response may contain an account page, locale chooser, or consent document. Check content type, final URL, a structural marker, and plausible record counts.

Emit structured diagnostics that help operations without retaining sensitive bodies: source host, status, duration, final URL class, matched containers, accepted records, and rejection reason. Keep logs separate from output records so a downstream consumer never mistakes an error page title for a scraped product.

Conclusion

Web scraping with Go is strongest when a small HTTP core feeds a deterministic parser and a bounded worker pool. Reuse one client, close every body, pass contexts through request boundaries, and make pagination state explicit. When the response HTML is incomplete because the page depends on JavaScript, swap the acquisition layer while preserving the goquery selectors and typed output.

The final acceptance test should be intentionally small: one known page, one expected identity marker, one plausible record range, and one serialized output sample that downstream code can read. That test catches transport, parsing, normalization, and schema mistakes without turning a production crawl into the first place the system proves itself.

Ready to Extend a Go Scraping Service?

Send rendered page content into the same net/http and goquery pipeline, then validate one public-data target with a bounded worker budget.

Sign up today and get $5 in free credit — no credit card required.

Claim Your $5 Credit →

FAQ

Is Go good for web scraping?

Go is a strong choice for long-running and concurrent scraping services because its standard HTTP package, explicit errors, and goroutines fit network-bound workloads. The main tradeoff is a smaller scraping ecosystem than Python or JavaScript.

Does goquery fetch web pages?

No. goquery parses and queries HTML; net/http or another client fetches the response. Keeping those jobs separate makes acquisition and selector tests easier to isolate.

How many goroutines should a Go scraper use?

A Go scraper should use a measured, host-aware limit rather than one goroutine per URL. Start with a small worker budget, observe target behavior and local memory, and increase only within authorized limits.

Why does a Go scraper return zero rows from a visible page?

A zero-row result often means the required nodes were created by JavaScript, the response is a different page, or the selector changed. Inspect the response body and identity markers before changing the selector.

Is web scraping with Go legal?

Web scraping with Go is subject to the same rules as any data collection method. Collect authorized public data, review site terms and robots directives, respect access controls, and consult counsel for regulated or consequential uses.

References