Back to Blog

Google Trends with Python: Collect Interest and Trending Topics

Alex Johnson
Alex Johnson

Senior Web Scraping Engineer

30-Sep-2026

TL;DR:

  • Google Trends interest is a normalized index, not a search count. Preserve the geography and time range needed to interpret each observation.
  • Trending Now RSS and historical interest answer different questions. The public feed provides current topics; it is not a replacement for an interest-over-time series.
  • Scrapeless exposes a dedicated google_trends tool. Discover the installed tool schema before choosing its query and data type.
  • Traffic buckets should remain text. A label such as a threshold must not become an exact count in a dashboard.
  • Free to start. New Scrapeless accounts include free Scraping Browser runtime — sign up at app.scrapeless.com.

Google Trends can tell you about relative search interest and about topics attracting current attention. Those are different data products. A report comparing a keyword's interest over time cannot substitute a list of today's trending topics, even if both contain the same phrase.

A Python collection workflow should choose the surface first, then preserve its settings. This guide connects a dedicated Trends tool for interest data and runs a separate public RSS workflow for current topics. It keeps the output boundaries visible instead of forcing incompatible metrics into one column.

If your application already uses structured actor workflows, treat Trends as its own actor/tool contract. Google Search API search results and Google Trends interest are separate datasets.

Google Trends interest-over-time values describe relative search interest within the selected geography and period. The series is scaled from 0 to 100, and low-volume terms can appear as zero without demonstrating that nobody searched for them.

Google Trends normalization and sampling explains why a point depends on the share of searches within the chosen conditions. A peak of 100 identifies a relative maximum in that series, not a known number of searches.

Changing the date range changes the comparison basis. A value from a one-month request is not automatically comparable to the same date in a longer request. Sampled data and low-volume suppression also limit the conclusions you can draw.

Surface Suitable question Keep with the result
Interest over time How did relative interest change? Query, geography, dates, category and time zone
Interest by subregion Where is relative interest concentrated? Geographic scope and normalization context
Related queries/topics Which searches or topics are associated? Data type and returned ranking/metric meaning
Trending Now RSS Which topics are trending now? Feed geography, publication text and observation time

The dedicated google_trends tool exposes Trends-specific query parameters rather than ordinary search-result parameters. It is available through the Scrapeless MCP connection and belongs to the broader Scraping API surface.

The locally discovered schema includes q, data_type, date, geo, hl, tz and cat. Examples of data_type are interest_over_time, interest_by_subregion, related_queries and related_topics. Select the type required by the report before mapping fields into storage.

A managed tool does not change the meaning of Google's metrics. The output still needs geography, sampling context and a distinction between missing data and zero values. Check Scrapeless pricing for current plan terms; there is no cost or freshness benchmark in this tutorial.

Prerequisites and package setup

Use Python 3.12, Node.js, MCP 2.2.0, Requests 2.34.2 and scrapeless-mcp-server 0.6.3. The RSS path needs only Requests and Python's XML/CSV libraries. Interest-data capture additionally needs a real SCRAPELESS_KEY with tool access.

Authenticated Trends collection remains pending live verification without that credential. Local tool discovery and the public RSS path run independently; neither is described as a completed authenticated interest-series capture.

Install the pinned clients and server in a disposable project:

bash Copy
python -m pip install mcp==2.2.0 requests==2.34.2
pnpm add scrapeless-mcp-server@0.6.3

Discover the interest-data contract before collecting

Tool discovery gives the client the exact schema exposed by its installed MCP server. A local startup value of metadata-discovery-only can be used only for schema inspection; it must never be used to invoke remote capture.

Save this script as trends_mcp.py. Run its default mode with that obvious discovery value, or with your configured real key. The bootstrap suppresses nonprotocol console logging so stdout stays usable for stdio messages. The Python client uses MCP 2.2.0's input_schema and is_error attributes.

python Copy
import argparse
import asyncio
import json
import os
from pathlib import Path
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client

async def main(capture):
    key = os.environ['SCRAPELESS_KEY']
    server = StdioServerParameters(
        command='node',
        args=['--input-type=module', '-e',
              'console.log = () => {}; await import(process.argv[1]);',
              str(Path('node_modules/scrapeless-mcp-server/build/index.js').resolve())],
        env={**os.environ, 'SCRAPELESS_KEY': key})
    async with stdio_client(server) as streams:
        async with ClientSession(*streams) as session:
            await session.initialize()
            listed = await session.list_tools()
            tool = next(t for t in listed.tools if t.name == 'google_trends')
            Path('trends-tool-schema.json').write_text(
                json.dumps(tool.input_schema, indent=2))
            print(json.dumps({'discovered_tools': len(listed.tools),
                              'selected_tool': tool.name}))
            if capture:
                if key == 'metadata-discovery-only':
                    raise ValueError('A real key is required for remote capture')
                args = {'q': 'coffee', 'data_type': 'interest_over_time',
                        'date': 'today 12-m', 'geo': 'US',
                        'hl': 'en', 'tz': '0', 'cat': '0'}
                result = await session.call_tool(tool.name, args)
                Path('trends-result.json').write_text(json.dumps({
                    'request': args, 'result': result.model_dump(mode='json')
                }, indent=2))
                if result.is_error:
                    raise ValueError('Tool error; inspect saved result')
                print('Raw result saved; inspect its actual fields before mapping a series.')

parser = argparse.ArgumentParser()
parser.add_argument('--capture', action='store_true')
asyncio.run(main(parser.parse_args().capture))

The local handshake discovered 25 tools and selected google_trends. That count describes the installed server, not an eternal platform limit. The saved trends-tool-schema.json is the evidence to inspect when the server package changes.

With a real key configured, python trends_mcp.py --capture requests interest over time for coffee in the US across the last year. Note: the capture branch requires credentials and remains pending authenticated live verification. The script preserves the MCP result envelope before making assumptions about its inner fields. It does not invent a timeline key or fabricate a time series.

Start Scraping with Scrapeless

Power up your web scraping and automation workflow with Scrapeless!
Sign up today and get $5 in free credit — no credit card required.

Claim your free credit now in the Scrapeless Dashboard.

Collect current topics from the public RSS feed

Trending Now RSS is a public current-topic feed that can be collected without a browser or Scrapeless API credential. Its publication timestamps and traffic labels describe the feed's topic observations, not historical normalized interest.

The following complete script makes one feed request and stores at most five items. It saves the original XML, JSON records and CSV in an observation directory. This is the independently executable public-data path.

python Copy
import csv
import hashlib
import json
from datetime import datetime, timezone
from pathlib import Path
from xml.etree import ElementTree as ET
import requests

url = 'https://trends.google.com/trending/rss?geo=US'
response = requests.get(url, timeout=30)
response.raise_for_status()
root = ET.fromstring(response.content)
channel = root.find('channel')
if channel is None:
    raise ValueError('Expected RSS channel, received another document')
observed = datetime.now(timezone.utc).strftime('%Y%m%dT%H%M%SZ')
folder = Path('trending-rss') / observed
folder.mkdir(parents=True, exist_ok=True)
(folder / 'source.xml').write_bytes(response.content)
ns = {'ht': 'https://trends.google.com/trending/rss'}
rows = []
for item in channel.findall('item')[:5]:
    title = (item.findtext('title') or '').strip()
    link = (item.findtext('link') or '').strip()
    if not title or not link:
        raise ValueError('RSS item missing title or link')
    rows.append({'title': title, 'link': link,
                 'published': item.findtext('pubDate'),
                 'traffic_bucket': item.findtext('ht:approx_traffic', namespaces=ns),
                 'geo': 'US', 'observed_at': observed})
if not rows:
    raise ValueError('No items; inspect source before treating this as valid empty data')
(folder / 'records.json').write_text(json.dumps(rows, ensure_ascii=False, indent=2))
with (folder / 'records.csv').open('w', newline='', encoding='utf-8') as out:
    writer = csv.DictWriter(out, fieldnames=list(rows[0]))
    writer.writeheader()
    writer.writerows(rows)
print(json.dumps({'source': response.url, 'items_saved': len(rows),
                  'source_sha256': hashlib.sha256(response.content).hexdigest()}))

The public feed run saved five actual items and their source hash. Topics change, so the article does not freeze the names into a supposed universal output sample. The record fields are title, link, published, traffic_bucket, geo and observed_at.

The ht namespace identifies the feed's optional traffic field. Keep its value as returned text, including any threshold notation. A missing traffic label is a missing field, not a zero-search estimate. The CSV writer follows the quoting and record rules described by the CSV interchange format, so topic titles containing punctuation remain intact.

A Trends storage table should retain the collection surface as well as the data. Give RSS topics and interest-series points separate metric names and separate acceptance rules.

For an interest series, keep the full request arguments with the raw result. Once authenticated output has been inspected, add a schema adapter that checks point types, query labels and the reported period. Do not convert a missing series into an empty successful chart simply because the HTTP or MCP exchange completed.

For RSS, retain the original feed bytes, feed geography and the observation directory. Publication text is source-supplied; observation time is when your client collected it. These are different clocks. The source hash connects each export to the feed from which it was derived, a practical use of provenance relationships.

A dashboard should show the selected date range beside the chart. Export that range with the points instead of keeping it only in a UI filter. Otherwise a spreadsheet recipient can interpret a normalized peak as a volume statistic that the data never supplied.

Common interpretation mistakes

A successful collection can still produce an incorrect analysis if the metric is mislabeled. The following checks belong at the point where collected data becomes a report.

Mistake Correct handling
Treating interest 100 as 100 searches Label the field as normalized interest
Comparing independently scaled date ranges Keep requests aligned or disclose the different basis
Treating an RSS traffic bucket as an exact count Preserve the original threshold text
Turning missing values into zero Preserve missing-state semantics and inspect the source
Mixing worldwide and country requests Group records by explicit geography
Calling general search for Trends history Use the dedicated Trends contract

Use related_queries or related_topics when you need associations, and keep the chosen data type in the record. Do not label all related results as current rising topics unless the returned dataset actually identifies them that way.

Conclusion: store the question with the metric

Google Trends with Python is most useful when the collection path matches the question being asked. Historical interest belongs to a Trends-specific contract; current-topic collection has a simple public RSS path.

Begin with the RSS export to establish evidence handling, then inspect the first authenticated interest result before building its adapter. Keep geography, date range and metric identity attached to every downstream chart.


Ready to Build Your AI-Powered Data Pipeline?

Join our community to claim a free plan and connect with developers building web-data pipelines: Discord · Telegram.

Sign up at app.scrapeless.com for free Scraping Browser runtime and adapt the patterns above to your own public-data workflow.


FAQ

Q: Does Google Trends report exact search volumes?

No. Google Trends interest values are normalized measures of relative interest. They cannot be converted into exact search counts from the index alone.

Q: Can Python collect Trending Now without an API key?

Yes. The public RSS example collects current topics with Requests and standard-library parsers. It does not retrieve a historical interest-over-time series.

Q: Is the Google Search API the same as the Trends tool?

No. Search-result collection and Google Trends collection have different request parameters and output meanings. Use google_trends for the Trends-specific workflow shown here.

Q: Is collecting public Google Trends data legal?

The acceptable scope depends on terms, access conditions, usage rights and jurisdiction. Review those constraints for the intended collection and reuse; a public feed is not unrestricted permission for every application.

Q: Does the RSS workflow need a proxy or browser?

No proxy or browser is used by this RSS example. If the response is not the expected RSS document, stop the export and inspect the captured response rather than accepting a challenge page as topic data.

Q: Can Trends values differ between collections?

Yes. Sampling, query interpretation, geography and time-range choices affect the series. Record those choices so a difference can be investigated rather than automatically treated as a demand change.

Q: Can this workflow run without an AI agent?

Yes. Both the direct Python MCP client and RSS parser are ordinary scripts. An AI agent is optional; it does not supply missing API credentials or replace metric validation.

At Scrapeless, we only access publicly available data while strictly complying with applicable laws, regulations, and website privacy policies. The content in this blog is for demonstration purposes only and does not involve any illegal or infringing activities. We make no guarantees and disclaim all liability for the use of information from this blog or third-party links. Before engaging in any scraping activities, consult your legal advisor and review the target website's terms of service or obtain the necessary permissions.

Most Popular Articles

Catalogue