Home / n8n Workflow Templates / Scrape Top GitHub Trending Repositories Daily

Scrape Top GitHub Trending Repositories Daily

Scrape Today's Github Trend 13 Top Repositories

Automatically scrapes and processes the top repositories from GitHub's trending page, extracting key details like author, title, description, and URL.

Manual7 nodesDeveloper Toolsgithubscraperdeveloper

About this workflow

Scrape Today's GitHub Trending Repositories Workflow Analysis

This n8n workflow template is designed to automatically scrape and extract structured data from the GitHub Trending page (https://github.com/trending), specifically targeting the top repositories featured for the current day. It leverages HTML parsing capabilities to navigate GitHub’s public HTML structure, isolate relevant repository elements, and transform raw scraped content into a clean, usable list of repository details including author, name, description, and direct URL.

What the Workflow Does

The workflow initiates manually, fetches the HTML content of GitHub’s trending page, isolates the main container holding repository listings, extracts individual repository cards, parses key metadata (such as repository path, programming language, and description), and finally formats each entry into a standardized JSON object with enriched fields like author, title, url, and timestamped created_at. The output is a list of up to 25 trending repositories (as displayed by GitHub), ready for downstream use such as database insertion, notification alerts, or integration with other tools.

Key Features and Capabilities

  • Web Scraping Without APIs: Bypasses the need for GitHub API tokens by directly scraping public HTML.
  • Dynamic Data Extraction: Uses CSS selectors to reliably target repository elements even as GitHub’s layout evolves moderately.
  • Structured Output: Converts unstructured HTML into consistent JSON records with normalized fields.
  • Manual Trigger: Designed for on-demand execution via the “Test workflow” button, ideal for scheduled or ad-hoc runs.
  • Text Cleaning: Includes options to trim whitespace and clean up extracted text for better readability and consistency.
  • URL Construction: Dynamically builds valid GitHub repository URLs from scraped paths.

Main Nodes and Their Purposes

  1. When clicking ‘Test workflow’ (manualTrigger)

    • Serves as the entry point; triggers the workflow manually during testing or execution.
  2. Request to Github Trend (httpRequest)

    • Sends an HTTP GET request to https://github.com/trending to retrieve the raw HTML of the trending page.
  3. Extract Box (html node)

    • Parses the full HTML response and extracts the content of the first <div class="Box">, which contains the main list of trending repositories.
  4. Extract all repositories (html node)

    • Processes the extracted "Box" HTML and isolates each <article class="Box-row"> element (representing individual repositories) into an array of HTML snippets using returnArray: true.
  5. Turn to a list (splitOut node)

    • Splits the array of repository HTML snippets into individual items, enabling item-by-item processing in subsequent nodes.
  6. Extract repository data (html node)

    • For each repository HTML snippet, extracts three key pieces of information:
      • repository: The author/repo path from the <a class="Link"> element.
      • language: Programming language from <span class="d-inline-block">.
      • description: Brief project description from the <p> tag.
  7. Set Result Variables (set node)

    • Transforms and enriches the extracted data:
      • Splits repository into author and title.
      • Constructs a valid url using the repository path.
      • Adds a created_at timestamp using {{$now}}.
      • Preserves the original description.
      • Retains all other fields via includeOtherFields: "=".

Use Cases and Benefits

  • Developer Research: Quickly discover new or rising open-source projects.
  • Automated Monitoring: Track daily trends for competitive analysis or inspiration.
  • Content Aggregation: Feed trending repos into dashboards, newsletters, or internal tools.
  • Educational Purposes: Learn web scraping techniques using n8n’s visual automation.
  • Low-Cost Alternative: Avoid rate limits or authentication requirements of the official GitHub API for public data.

Step-by-Step Workflow Logic

  1. Trigger: User clicks “Test workflow,” activating the manual trigger.
  2. Fetch Page: The httpRequest node retrieves the full HTML of https://github.com/trending.
  3. Isolate Container: The Extract Box node uses div.Box to capture the primary trending section.
  4. Extract Repository Cards: The Extract all repositories node selects all article.Box-row elements and returns them as an array of HTML strings under the repositories field.
  5. Split into Items: The Turn to a list node converts the repositories array into individual workflow items for parallel or sequential processing.
  6. Parse Metadata: For each repository item, the Extract repository data node applies CSS selectors to pull out:
    • Repository path (e.g., "n8n-io/n8n")
    • Language badge text
    • Description paragraph
  7. Enrich & Standardize: The Set Result Variables node:
    • Splits the repository path into author and title
    • Builds a clickable GitHub URL
    • Adds current timestamp
    • Cleans and assigns all fields into a uniform structure
  8. Output: The final output is a list of structured repository objects, ready for export, storage, or further automation.

This workflow exemplifies efficient, no-code web scraping using n8n’s native HTML parsing and data manipulation nodes, providing actionable insights from publicly available GitHub data.

How to use: download the JSON, then in n8n choose “Import from File”.