Home / n8n Workflow Templates / Automated Document Embedding and Vector Storage with OpenAI and Qdrant

Automated Document Embedding and Vector Storage with OpenAI and Qdrant

Qdrant Vector Database Embedding Pipeline

This workflow automates the process of fetching JSON files via FTP, converting them into documents, splitting text into chunks, generating OpenAI embeddings, and storing them in a Qdrant vector database for semantic search.

Manual9 nodesAI & MLAIVector DatabaseData Processing

About this workflow

Qdrant Vector Database Embedding Pipeline: A Comprehensive Overview

This N8n workflow template, titled "Qdrant Vector Database Embedding Pipeline", automates the process of extracting, processing, and embedding JSON documents into a vector database (Qdrant) using OpenAI's embeddings. It is designed for semantic search applications where unstructured or semi-structured textual data must be transformed into searchable, high-dimensional vectors.

The pipeline integrates file handling via FTP, document parsing, text chunking, AI-powered embedding generation, and vector storage—all orchestrated within the n8n automation platform. This enables organizations to build scalable retrieval-augmented generation (RAG) systems, knowledge bases, or intelligent search engines with minimal manual intervention.


What the Workflow Does

This workflow performs an end-to-end data ingestion and vectorization pipeline:

  1. Connects to an FTP server and lists all .json files in a specified directory.
  2. Iteratively downloads each file in binary format.
  3. Parses the JSON content into LangChain-compatible document objects.
  4. Splits the text content into smaller chunks based on a defined separator ("chunk_id").
  5. Generates embeddings for each text chunk using OpenAI’s text-embedding-ada-002 model.
  6. Stores the embedded chunks—including their metadata—into a Qdrant vector database collection named sv_lang_data.

It supports batched processing and ensures efficient handling of large datasets by leveraging streaming and batching mechanisms.


Key Features and Capabilities

  • Automated File Ingestion: Uses FTP nodes to discover and download JSON files automatically.
  • Scalable Batching: Employs Split In Batches node to process one file at a time, enabling control over memory usage and API rate limits.
  • Flexible Document Parsing: Utilizes the Default Data Loader to convert binary JSON payloads into structured LangChain documents.
  • Configurable Text Chunking: Applies character-based splitting using "chunk_id" as a delimiter, allowing logical segmentation aligned with source data structure.
  • OpenAI-Powered Embeddings: Integrates securely with OpenAI API to generate 1536-dimensional embeddings optimized for semantic similarity.
  • Vector Storage in Qdrant: Pushes embedded documents into a pre-configured Qdrant collection, supporting fast approximate nearest neighbor (ANN) searches.
  • Metadata Preservation: Retains relevant metadata from original files during transformation and storage.
  • Visual Documentation: Includes sticky notes that explain each stage, improving readability and maintainability.

Main Nodes and Their Purposes

Node Name Type Purpose
When clicking ‘Test workflow’ Manual Trigger Initiates the workflow manually for testing purposes.
List all the files FTP Node (List Operation) Retrieves a list of all files in the remote FTP directory: Oracle/AI/embedding/svenska.
Loop over one item Split In Batches Processes one file entry at a time from the list, enabling sequential and controlled execution.
Downloading item FTP Node (Download) Downloads the current file (referenced via {{ $json.name }}) in binary format and stores it under binary.data.
Default Data Loader LangChain Document Loader Converts the downloaded binary JSON data into LangChain Document objects suitable for downstream NLP tasks.
Character Text Splitter LangChain Text Splitter Breaks down long texts into smaller chunks using "chunk_id" as the separator; prepares input for embedding models.
Embeddings OpenAI LangChain Embeddings Node Calls OpenAI API to generate vector embeddings for each text chunk (uses text-embedding-ada-002).
Qdrant Vector Store LangChain Vector Store Node Inserts the generated embeddings along with associated metadata into the sv_lang_data collection in Qdrant.

Additionally:

  • Sticky Notes provide inline documentation explaining each phase.
  • The Qdrant API credential ("QdrantApi svenska") and OpenAI API key are securely referenced through stored credentials.
  • The FTP account is reused across multiple nodes for listing and downloading operations.

Use Cases and Benefits

Use Cases

  • Enterprise Knowledge Management: Index internal documentation, support tickets, or training materials for semantic search.
  • Multilingual Semantic Search: Specifically targets Swedish language data (implied by "svenska"), making it ideal for Scandinavian markets.
  • Data Lake Integration: Automate ingestion from legacy systems storing data in JSON format on FTP servers.
  • RAG Pipelines: Feed processed and embedded content into LLM applications requiring contextual grounding.
  • Compliance & Archival Systems: Enable fast retrieval of archived records using natural language queries.

Benefits

  • End-to-End Automation: Eliminates manual steps in data preparation and embedding.
  • High Accuracy: Preserves context through intelligent chunking and uses state-of-the-art embeddings.
  • Efficient Scaling: Batch processing prevents overload; ideal for hundreds or thousands of files.
  • Secure Credential Handling: All sensitive keys (OpenAI, Qdrant, FTP) are managed via n8n’s encrypted credential system.
  • Extensible Design: Easy to modify for other languages, embedding models, or alternative vector databases (e.g., Pinecone, Weaviate).

Step-by-Step Workflow Logic

  1. Trigger Execution

    • The workflow starts when a user clicks “Test workflow” (manual trigger).
    • Control passes to the List all the files node.
  2. Fetch File List from FTP

    • The FTP node connects to the server and lists all files in /Oracle/AI/embedding/svenska.
    • Outputs an array of file entries (name, size, date, etc.).
  3. Iterate Over Files One-by-One

    • The Split In Batches node processes one file per iteration.
    • Ensures stability and avoids overwhelming downstream services.
  4. Download Current File

    • Another FTP node downloads the selected file using dynamic path:
      =Oracle/AI/embedding/svenska/{{ $json.name }}
    • Saves the file as binary data in binary.data.
  5. Parse Binary JSON into Documents

    • The Default Data Loader reads the binary JSON and converts it into LangChain Document objects.
    • Each document contains pageContent (text) and metadata.
  6. Split Text into Chunks

    • The Character Text Splitter divides the document text wherever "chunk_id" appears.
    • Ensures semantic boundaries are respected (e.g., per-section or per-record splits).
    • Optional but recommended for uniform vectorization.
  7. Generate Embeddings

    • The Embeddings OpenAI node sends each text chunk to OpenAI’s API.
    • Returns a 1536-dimensional vector representing the semantic meaning of the text.
  8. Store Vectors in Qdrant

    • The Qdrant Vector Store node receives both embeddings and original documents.
    • Inserts them into the sv_lang_data collection.
    • Uses a batch size of 100 for optimal performance.
    • Assumes Qdrant collection is pre-created with:
      {
        "vectors": {
          "size": 1536,
          "distance": "Cosine"
        }
      }
      
  9. Repeat Until Complete

    • The loop continues until all files have been processed.
    • Final output confirms successful indexing of all documents.

This workflow exemplifies how n8n can orchestrate complex AI/data pipelines without writing code, combining file operations, language models, and vector databases into a robust, production-ready solution.

How to use: download the JSON, then in n8n choose “Import from File”.