Semantic Chunking for RAG: Build a Web Data Pipeline
Advanced Data Extraction Specialist
TL;DR:
- Semantic chunking groups text using changes in meaning. It needs a defined segmentation method, an embedding model, and a rule for deciding where a chunk ends.
- Clean source text is part of retrieval quality. Navigation menus and repeated banners can distort the chunks before the retriever sees them.
- Every chunk should retain its source and position. Provenance makes a retrieved passage useful for citations and diagnosis.
- A semantic splitter needs a baseline comparison. Extra embedding work does not guarantee better answers than a simpler split strategy.
- Free to start. Use Scrapeless account credit to evaluate the upstream collection step on a small approved document set.
Introduction: Retrieval Starts Before the Vector Store
A retriever can only return the passages stored in its index. In retrieval-augmented generation, the retrieved evidence forms part of the generation input. If a passage contains half an explanation, or combines a policy statement with unrelated navigation text, the answer generator receives an incomplete context even when the similarity score looks strong.
Semantic chunking changes where those passages begin and end. Instead of cutting only at a fixed length, the splitter compares neighboring sentence representations and creates a boundary when their meanings diverge enough. A size limit still matters because a coherent topic can be longer than the downstream model can accept.
This article describes a web-to-RAG pipeline with Scrapeless as the acquisition layer and a transparent Python splitter as the transformation layer. It complements structured extraction with a language model: extraction produces task fields, while chunking prepares passages for later retrieval.
What Semantic Chunking Does
Semantic chunking uses a representation of meaning to choose text boundaries. A common implementation splits a document into sentences, embeds those sentences, compares adjacent vectors, and starts a new group at a sufficiently large change.
Sentence embedding methods map sentence text into vectors that can be compared for similarity. The meaning of a similarity score depends on the model and the input, so a numerical boundary setting is a parameter to evaluate rather than a universal definition of topic change.
| Strategy | Boundary rule | Useful starting point | Main trade-off |
|---|---|---|---|
| Fixed size | A configured length | A simple baseline | Can cut through an explanation |
| Structure aware | Headings and paragraph breaks | Organized reference pages | Depends on clean document structure |
| Semantic | Changes in embedding similarity | Text with topic shifts | Adds model computation and tuning |
| Hybrid | Structure, semantic checks, and limits | Mixed document collections | More configuration to evaluate |
Headings, tables, code blocks, and lists need special treatment. A sentence splitter can damage a table by detaching a value from its column label. Preserve these units during preprocessing or route them to a structure-aware path.
Pipeline at a Glance
The pipeline converts a captured web document into indexed passages that retain their origin. Its stages are acquisition → text normalization → sentence segmentation → embedding → chunk formation → retrieval evaluation.
Scrapeless Web Unlocker handles the target request. Your application cleans the returned content, generates embeddings through a selected model, and stores the resulting passages. Web acquisition does not itself create a vector index or choose the correct chunking strategy.
Save the original response before transformation. Keep the requested URL, any established canonical URL, collection time, page title, and a hash of the normalized text. A canonical URL must come from source evidence; do not manufacture it by deleting path segments.
Prerequisites
You need an approved set of public source pages, a Scrapeless API key for acquisition, Python for processing, and an embedding model you can run locally or through your chosen provider. The provider or model determines the embedding setup and access requirements.
This tutorial's splitter consumes a sentences.json file with a source_url string, an ordered sentences list, and an embeddings list containing one numeric vector per sentence. It does not assume a particular hosted model or vector dimension. The same model must be used for every sentence in one run.
Authenticated page acquisition and embedding generation are prerequisites that have not been executed end to end for this example. The splitter can be checked independently, but local algorithm checks do not establish retrieval quality or successful model integration.
Stage 1: Capture and Clean the Source Document
A useful document starts with the intended page content rather than a challenge page, login form, or navigation shell. Acquire the page through the Web Unlocker request interface and verify that its expected content is present before extracting text.
For each accepted page, remove scripts, styles, repeated navigation, and unrelated promotional blocks. Preserve headings and paragraph boundaries. If the page is rendered dynamically, use the documented rendering configuration and check that the returned content includes the relevant text.
Store raw and normalized content separately. That makes it possible to diagnose whether a bad retrieved passage came from collection, cleanup, segmentation, or the embedding model. The provenance relationships between source and derived data are useful here: the chunk is derived from a specific version of the normalized document.
Do not silently combine the body of one page with a title or publication date from another. If a field is unavailable, keep it unset. Page collection time and source publication time are different metadata fields.
Stage 2: Segment Sentences and Generate Consistent Embeddings
Sentence segmentation produces the ordered text units that the semantic splitter will compare. Use a language-appropriate segmenter and inspect its treatment of abbreviations, decimals, headings, and bullet lists on your sample documents.
Avoid claiming that a basic period split works for every page. A paragraph containing an abbreviation or a code example can break into misleading fragments. For short reference pages, heading and paragraph boundaries may already provide sufficient structure without a semantic pass.
Generate one embedding for each sentence and preserve the exact order. Record the model identifier, normalization settings, and source-text hash alongside the vectors. Reject missing vectors, inconsistent dimensions, or non-finite values before computing similarities.
A paragraph-level context window can improve some sentence representations, but it changes the experiment. Keep that setting fixed when comparing split rules so an embedding-input change does not get misattributed to the boundary algorithm.
Start Scraping with Scrapeless
Power up your web scraping and automation workflow with Scrapeless!
Sign up today and get $5 in free credit — no credit card required.Claim your free credit now in the Scrapeless Dashboard.
Stage 3: Create Bounded Semantic Chunks
The splitter below starts a new chunk when adjacent sentence similarity drops below a configured threshold or the character budget would be exceeded. Save it as chunk_sentences.py.
Note: This script requires a real
sentences.jsonfile generated from your captured text and embedding model. Embedding generation remains a prerequisite; the example does not include a fabricated vector set or claim a measured retrieval improvement.
python
import json
import math
import os
from pathlib import Path
source = json.loads(Path("sentences.json").read_text())
sentences = source["sentences"]
vectors = source["embeddings"]
threshold = float(os.environ["SIMILARITY_THRESHOLD"])
max_chars = int(os.environ["MAX_CHUNK_CHARS"])
if not -1 <= threshold <= 1 or max_chars <= 0:
raise ValueError("Invalid chunking configuration")
if not sentences or len(sentences) != len(vectors):
raise ValueError("Each sentence needs one embedding")
dimension = len(vectors[0])
if not dimension:
raise ValueError("Embeddings must not be empty")
for text, vector in zip(sentences, vectors):
if not isinstance(text, str) or not text.strip():
raise ValueError("Sentence text must be nonempty")
if len(text) > max_chars:
raise ValueError("Segment oversized sentences before chunking")
if len(vector) != dimension or not all(math.isfinite(v) for v in vector):
raise ValueError("Invalid embedding dimension or value")
if sum(v * v for v in vector) == 0:
raise ValueError("A zero vector has no cosine direction")
def cosine(a, b):
numerator = sum(x * y for x, y in zip(a, b))
denominator = math.sqrt(sum(x*x for x in a) * sum(y*y for y in b))
return max(-1.0, min(1.0, numerator / denominator))
chunks, start, current = [], 0, []
for index, sentence in enumerate(sentences):
topic_change = index > 0 and cosine(vectors[index-1], vectors[index]) < threshold
too_long = len(" ".join(current + [sentence])) > max_chars
if current and (topic_change or too_long):
chunks.append({"start_sentence": start, "end_sentence": index,
"text": " ".join(current)})
current, start = [], index
current.append(sentence)
chunks.append({"start_sentence": start, "end_sentence": len(sentences),
"text": " ".join(current)})
records = [dict(chunk, source_url=source["source_url"], chunk_index=index)
for index, chunk in enumerate(chunks)]
Path("chunks.json").write_text(json.dumps(records, ensure_ascii=False, indent=2))
print(json.dumps({"chunk_count": len(records)}))
The sentence interval uses an inclusive start and exclusive end. That convention makes it possible to reconstruct exactly which input sentences produced a chunk. Empty inputs and invalid vectors fail explicitly instead of producing an apparently successful empty index.
The size limit here counts characters. It is not a token limit. If your embedding or generation system has a token budget, measure tokens with that system's tokenizer before sending text. The code rejects an oversized sentence so the upstream segmenter can handle it deliberately.
This implementation is intentionally simple: it compares neighboring sentence vectors. It does not implement clustering across distant parts of a document, adaptive percentile thresholds, or a trained boundary classifier. Those are separate approaches that need their own evaluation.
Stage 4: Index Passages Without Losing Their Source
Indexing should preserve enough metadata to reconnect a retrieved chunk to its source document. Store the chunk text, source URL, text hash, chunk index, sentence range, collection time, and embedding model identifier.
Create chunk embeddings from the final chunk text if that is the representation your retriever uses. Averaging sentence embeddings is a different design choice; do not assume it produces the same retrieval ranking as embedding the assembled passage.
When a source document changes, keep the old and new versions distinguishable. Reusing the same chunk identifier while changing the underlying text can make stored citations ambiguous. A version key based on the normalized source hash helps separate observations.
A vector store is only one possible index. A small evaluation can retrieve passages in memory. Start with the simplest arrangement that allows you to inspect the text returned for each question.
Stage 5: Compare Retrieval, Not Just Chunk Appearance
Semantic chunking should be evaluated on the questions the application needs to answer. A visually tidy set of passages does not prove that the retriever will return the right evidence.
Build a small question set with supporting passages marked in the original documents. Run a fixed-size baseline, a structure-aware baseline, and the semantic splitter against the same normalized content. Hold the embedding model, retrieval settings, and answer evaluation constant.
| Measurement | Question it answers |
|---|---|
| Supporting-passage retrieval | Was the necessary evidence retrieved? |
| Unrelated retrieved text | How much context was wasted? |
| Answer support | Does the generated answer follow the evidence? |
| Processing time | What did segmentation and embedding cost operationally? |
| Input volume | How much text reached each model step? |
Research on semantic chunking's computational trade-offs supports treating the method as an empirical choice. The extra work does not consistently produce an improvement across every task and dataset.
Inspect failures by stage. Missing evidence in the raw capture is an acquisition problem. Broken paragraphs are a cleanup problem. A relevant passage ranked below irrelevant passages may point to the embedding or retrieval configuration. These problems require different changes.
Track the acquisition component separately through Scrapeless pricing. Embedding and indexing costs belong to the downstream stack; they should not be reported as Web Unlocker usage.
Conclusion: Choose the Splitter That Retrieves Better Evidence
Semantic chunking is useful when meaning changes inside documents in ways that simpler boundaries fail to capture. The practical workflow is to preserve clean source text, produce consistent embeddings, enforce passage limits, and compare retrieval against simpler alternatives.
Keep the source metadata attached throughout. A retrieved passage becomes much more useful when the application can show where it came from and which document version it represents.
Ready to Build Your Web Data Pipeline?
Join developers working on web data collection in Discord and Telegram.
Create a Scrapeless account and adapt the workflow to your own approved data sources.
FAQ
Q: Is semantic chunking always better than fixed-size chunking?
No. Its benefit depends on the documents, embedding model, retrieval settings, and task. Compare strategies on a shared evaluation set before choosing one.
Q: Does Scrapeless generate the embeddings in this workflow?
No. Scrapeless supplies the upstream web-access step. Your selected embedding model and index handle the downstream representation and retrieval.
Q: How should tables and code blocks be chunked?
Preserve their structure or route them through a dedicated structure-aware splitter. Sentence boundaries alone can detach values from headers or break code context.
Q: What if an acquired page contains a challenge or different HTML?
Reject content that fails the source contract and inspect the acquisition or parser stage. Semantic similarity cannot repair an incorrect source document.
Q: Does collection need a separate proxy?
Use the routing options of your chosen Web Unlocker request. Do not infer proxy requirements from the chunking algorithm, which runs after collection.
Q: Can the pipeline run without an AI agent?
Yes. An application can acquire documents and run embedding-based chunking without an autonomous agent. An embedding model is still required for the semantic comparison.
Q: What limits should apply to parallel source collection?
Start with bounded collection, such as no more than three workers per host, and honor stricter target and account limits. Embedding batches have separate model-specific limits.
Q: Can any public page be added to a RAG index?
Public availability alone does not establish permission for every collection or reuse. Review source terms, applicable rights, and the intended data handling before ingestion.
At Scrapeless, we only access publicly available data while strictly complying with applicable laws, regulations, and website privacy policies. The content in this blog is for demonstration purposes only and does not involve any illegal or infringing activities. We make no guarantees and disclaim all liability for the use of information from this blog or third-party links. Before engaging in any scraping activities, consult your legal advisor and review the target website's terms of service or obtain the necessary permissions.



