Skip to main content

Data Flow

Detailed sequence diagrams showing how data flows through Rephole.

Repository Ingestion Flow​

When you submit a repository for ingestion:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Client │────▢│ API Server│────▢│ Redis Queue │────▢│ Worker β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”˜
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ PostgreSQL │◀────│ ChromaDB │◀────│ Clone β†’ Parse β†’ Embedβ”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Step by Step​

  1. Client sends POST /ingestions/repository
  2. API Server validates request and creates job
  3. Redis Queue stores job for processing
  4. API Server returns jobId immediately
  5. Worker picks up job from queue
  6. Worker clones repository to local storage
  7. Worker parses code files using Tree-sitter
  8. Worker generates embeddings via OpenAI API
  9. Worker stores vectors in ChromaDB
  10. Worker stores file content in PostgreSQL
  11. Worker marks job complete in Redis

Semantic Search Flow​

When you perform a search:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Client │────▢│ API Server│────▢│ OpenAI API β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β”‚ β”‚
β”‚β—€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚ (embedding)
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ ChromaDB β”‚
β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
β”‚ (child chunk IDs)
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ PostgreSQL β”‚
β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
β”‚ (parent content)
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Response β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Step by Step​

  1. Client sends POST /queries/search with prompt
  2. API Server sends prompt to OpenAI for embedding
  3. OpenAI returns 1536-dimensional vector
  4. API Server queries ChromaDB for similar child chunks
  5. ChromaDB returns matching chunk IDs (k Γ— 3)
  6. API Server groups chunks by parent document
  7. API Server fetches parent content from PostgreSQL
  8. API Server formats and returns results to client

Data Storage​

ChromaDB (Vectors)​

Collection: rephole-collection
β”œβ”€β”€ Document ID: chunk-001
β”‚ β”œβ”€β”€ Vector: [0.123, -0.456, ...] (1536 dims)
β”‚ └── Metadata: {parent_id, file_path, repo_id}
β”œβ”€β”€ Document ID: chunk-002
β”‚ └── ...

PostgreSQL (Content)​

Table: files
β”œβ”€β”€ id: ulid
β”œβ”€β”€ repo_id: string
β”œβ”€β”€ path: string
β”œβ”€β”€ content: text (full file)
β”œβ”€β”€ created_at: timestamp
└── updated_at: timestamp

Queue Structure​

BullMQ Job​

{
"name": "repo-ingestion",
"data": {
"repoUrl": "https://github.com/...",
"ref": "main",
"token": "...",
"userId": "...",
"repoId": "..."
},
"opts": {
"attempts": 3,
"backoff": {
"type": "exponential",
"delay": 1000
}
}
}