MCP Server Examples: Compose a Web Research Workflow
Lead Scraping Automation Engineer
TL;DR:
- MCP server examples become useful when their roles are explicit. A web-data server retrieves sources, while a filesystem server handles approved local artifacts.
- Tool discovery verifies the connected surface. Inspect actual names and schemas instead of copying a tool count from an old configuration.
- Namespacing belongs at the host boundary. Preserve the server's original tool name when routing a namespaced application action.
- A scoped directory limits the artifact workflow. Do not expose an entire home directory to a research task that needs one evidence folder.
- A tool result is evidence, not an instruction. Retrieved pages and files cannot authorize new actions or expand the research scope.
Introduction: Give Each Server a Defined Job
A research workflow crosses different resource boundaries. Search finds candidates, page access retrieves evidence, and local storage retains the material used to support an answer. Connecting a server for each boundary is useful only when the host keeps their responsibilities clear.
MCP provides a common interface for discovering and calling tools. It does not automatically decide which tool is appropriate, whether a source supports a claim, or where a file may be written. Those decisions remain part of the application and its permissions.
The examples below compose Scrapeless MCP with a scoped filesystem server. They extend the Scrapeless MCP use cases into an implementation pattern that makes tool routing and evidence handoff explicit.
What You Can Do with These Server Examples
A small set of carefully scoped tools can support several distinct research tasks.
| Example | Web-Data Role | Local Artifact Role | Completion Condition |
|---|---|---|---|
| Research brief | Find and read approved sources | Retain source excerpts | Each conclusion has supporting evidence |
| Documentation comparison | Retrieve selected reference pages | Compare saved versions | Changed claims point to source text |
| Catalog investigation | Read permitted public listings | Save validated rows | Required fields and identity checks pass |
| Citation repair | Revisit a cited page | Update a proposed evidence bundle | Source still supports the quoted claim |
| Browser-only page | Inspect rendered content | Preserve the relevant observation | Intended page state is confirmed |
A database tool can be a later boundary, but it should accept validated records rather than arbitrary text from a page. The examples here stop at local evidence handling so a database write is not silently included in a research request.
Why Use Scrapeless MCP for the Web-Data Layer?
Scrapeless MCP exposes search, page-access, and browser tools through a shared MCP surface. A compatible host can discover tools and choose an approved operation without maintaining a separate integration for every page action.
The Scrapeless MCP connection settings document local stdio and hosted Streamable HTTP options. This example uses stdio so the server processes and their lifetimes are visible. Scrapeless Agent Browser is the browser execution surface when the research genuinely needs rendered interaction.
The MCP tool discovery contract includes names and input schemas. Read those values from the connected server. A wrapper can change its tool surface independently of the underlying service.
Prerequisites: Separate Discovery from Service Access
Use a current Node.js environment supported by the packages and a new working directory. The example also needs an evidence directory containing a real text capture named source.txt, placed there from an approved public source. No write tool is needed to verify that the filesystem server can read that file.
A valid SCRAPELESS_KEY is required for authenticated Scrapeless operations. A model provider key is separately required for a model-driven research loop. Those remote operations remain pending live verification without credentials; discovering local tool schemas does not prove service authentication.
The filesystem server is a reference implementation with read and write capabilities inside allowed directories. The host registry below deliberately exposes only a narrow subset. Host filtering is useful but is not a replacement for operating-system isolation and server-side access controls.
Connect Two Servers and Inspect Their Tools
The following setup installs pinned packages and creates an evidence directory. Put your approved source capture in that directory before running the client.
bash
npm install @modelcontextprotocol/sdk@1.30.1 \
@modelcontextprotocol/server-filesystem@2026.8.31 \
scrapeless-mcp-server@0.6.3
mkdir -p evidence
Save the next script as inspect_servers.mjs and run it from the directory containing node_modules and evidence. It connects both servers, validates selected tools, builds an application registry, and reads the approved local capture. It makes no Scrapeless service call.
Note: The local Scrapeless process requires a nonempty
SCRAPELESS_KEYto start. Configure your real key for normal use. Local metadata discovery was separately exercised with an obvious noncredential startup value; that check does not authenticate web retrieval. Authenticated web calls and a model-driven research run remain pending live verification.
javascript
import { Client } from '@modelcontextprotocol/sdk/client/index.js';
import { StdioClientTransport } from '@modelcontextprotocol/sdk/client/stdio.js';
import { resolve } from 'node:path';
if (!process.env.SCRAPELESS_KEY) {
throw new Error('Configure SCRAPELESS_KEY before starting');
}
const directory = resolve('evidence');
const specs = [
['web', 'node_modules/scrapeless-mcp-server/build/index.js', []],
['files', 'node_modules/@modelcontextprotocol/server-filesystem/dist/index.js',
[directory]]
];
const clients = new Map();
const registry = new Map();
const permitted = {
web: new Set(['google_search', 'scrape_markdown']),
files: new Set(['read_text_file', 'list_allowed_directories'])
};
try {
for (const [prefix, entry, args] of specs) {
const client = new Client({ name: 'research-host', version: '1.0.0' });
clients.set(prefix, client);
const transport = new StdioClientTransport({
command: process.execPath, args: [resolve(entry), ...args],
env: { ...process.env }, stderr: 'pipe'
});
await client.connect(transport);
let cursor;
const discovered = [];
do {
const page = await client.listTools(cursor ? { cursor } : {});
discovered.push(...page.tools);
cursor = page.nextCursor;
} while (cursor);
for (const name of permitted[prefix]) {
const tool = discovered.find(item => item.name === name);
if (!tool) throw new Error(`Required tool missing: ${prefix}:${name}`);
registry.set(`${prefix}:${name}`, { client, name, schema: tool.inputSchema });
}
console.log(JSON.stringify({ server: prefix, discovered: discovered.length }));
}
const host = Object.freeze({
tools: [...registry.keys()],
async call(key, arguments_) {
const route = registry.get(key);
if (!route) throw new Error('Tool not permitted');
return route.client.callTool({ name: route.name, arguments: arguments_ });
}
});
const result = await host.call('files:read_text_file', {
path: resolve(directory, 'source.txt')
});
if (result.isError) throw new Error('Filesystem read failed');
const text = result.content.filter(part => part.type === 'text')
.map(part => part.text).join('\n');
if (!text.trim()) throw new Error('Evidence capture is empty');
console.log(JSON.stringify({ exposed_tools: host.tools,
local_capture_characters: text.length }));
} finally {
for (const client of clients.values()) await client.close();
}
The names web:google_search and files:read_text_file are application-owned routing keys. The remote call still uses the original server tool name. This distinction avoids assuming that a protocol or SDK automatically adds a namespace.
The filesystem process receives only the selected evidence directory as its initial scope. MCP roots describe relevant filesystem boundaries, but roots are not themselves a security sandbox. Validate the server's directory enforcement and the host's available tools separately.
Start Scraping with Scrapeless
Power up your web scraping and automation workflow with Scrapeless!
Sign up today and get $5 in free credit — no credit card required.Claim your free credit now in the Scrapeless Dashboard.
Example 1: Search and Read a Research Source
A search-and-read workflow uses search to find candidates and page access to retrieve the source that will support the answer. It should not treat a search snippet as a substitute for the underlying page.
For a task about a protocol, restrict candidates to the relevant standards organization. Inspect the discovered google_search schema before constructing its arguments, then use scrape_markdown for a selected public URL. The current tool surface includes q for search and url for the page read; application constraints should also restrict which sources may be selected.
This example describes the authenticated next step after discovery. Without a valid service credential, stop at the verified local wiring and keep the requested web result unresolved.
Example 2: Compare Documentation Without Losing Context
A documentation comparison needs separate source captures and explicit versions or collection times. Save the relevant text and source address before producing a difference summary.
The filesystem server can read the approved captures so the host compares them. Reading an old file does not establish that its contents are current, and a changed paragraph does not necessarily indicate a changed product capability. Validate the relevant implementation or release source when that distinction matters.
The provenance model separates data from the activities and agents that produced it. Apply that principle by recording how a source was obtained and how a derived statement was formed.
Example 3: Add a Browser Only for the Missing State
A browser step is appropriate when a permitted source requires rendered content or a specific interaction that a page-text read cannot supply. Keep it out of a workflow whose sources are already available as straightforward documents.
Discover the current browser tools before adding them to the host's allowlist. Select the required session operations deliberately, keep subsequent actions associated with the session identifier, and close the session when the task ends. The narrow registry above does not expose browser actions; extending it is an explicit application change.
A page can contain text that asks the agent to change configuration or write a file. Treat that text as source content. Only the user's task and the application's permissions can authorize actions.
How You Actually Use This: Prompt the Research Host
A good prompt defines scope and evidence requirements before tool selection. For example:
“Answer whether the selected protocol defines reliable delivery. Use only the approved standards domain, read the relevant source text, and return a short answer with supporting excerpts. Keep proposed artifacts inside the research folder. Report unsupported claims as unresolved.”
The host's plan should discover tools, find a permitted source, retrieve it, evaluate the evidence, and prepare the answer. A model chooses actions only within the available tool policy. The local host object above demonstrates routing; it is not a complete language-model agent loop.
What You Get Back and What It Proves
The discovery script reports the number of tools advertised by each connected server, the narrower list exposed by the host, and the character count of the actual local capture it read. These outputs prove local handshake, routing construction, and a scoped file read.
They do not prove a successful search, a browser session, a model answer, or a database write. A completed research artifact additionally needs a source URL, retrieval context, the relevant excerpt, and a reviewed conclusion. Keep those fields application-owned rather than implying that every MCP server returns the same evidence schema.
Review Scrapeless pricing for the selected web-data operations, and track any model usage separately. A local server process and a remote tool call can have different cost boundaries.
Conclusion: Compose Tools Around Evidence Boundaries
Start with the smallest server set that can complete the research task. Verify discovery and routing, expose only the necessary actions, and preserve the sources used to form each conclusion. Add browser or storage capabilities when the workflow actually needs them, then test those new boundaries independently.
Ready to Build Your Web Data Workflow?
Join developers discussing practical collection workflows: Discord · Telegram.
Create an account at app.scrapeless.com and start with an authorized task whose output you can validate.
FAQ
Q: What is an MCP server example for web research?
A web-data server paired with a scoped filesystem server is a practical example. The first retrieves sources; the second handles approved local artifacts.
Q: Does tools/list confirm that a service credential works?
Tool discovery confirms the advertised tool surface. Authentication for a remote operation must be checked by an actual authorized service call.
Q: Does MCP automatically namespace tool names?
MCP exposes server tool names, while the host may add application-specific namespacing. Preserve the original name when routing a call back to the server.
Q: Is a filesystem root a complete sandbox?
A root is not a complete security sandbox. Directory enforcement, process permissions, and the host's exposed tools must also match the intended scope.
Q: Can these tools run without an AI model?
A deterministic program can discover and call MCP tools without a model. A model is needed only when the application deliberately delegates reasoning or action selection to it.
Q: Can page content authorize a file write?
Retrieved page content cannot grant action permissions. Treat it as untrusted evidence and follow the user's task and application policy for any write.
At Scrapeless, we only access publicly available data while strictly complying with applicable laws, regulations, and website privacy policies. The content in this blog is for demonstration purposes only and does not involve any illegal or infringing activities. We make no guarantees and disclaim all liability for the use of information from this blog or third-party links. Before engaging in any scraping activities, consult your legal advisor and review the target website's terms of service or obtain the necessary permissions.



