- “How did our gross margin trend across the last four quarters?”
- “What did the CFO say about guidance in Q2 vs Q4?”
- “Which metric deteriorated most between Q3 2023 and Q1 2024?”
- “Summarise all references to churn risk across our last six board memos.”
recency_bias to favor recent material.
API setup: Base URL: https://api.hydradb.com. Get your API key at app.hydradb.com.
Prerequisites
Required knowledge: Python basics, REST APIs, environment variablesRequired tools:
- HydraDB API key
- Python 3.11 or 3.12 (
python --version) pip install hydradb-sdk- An OpenAI API key and
pip install openai, for the steps that write answers from the retrieved chunks
What You’ll Build
By the end of this cookbook, you’ll be able to:- Ingest earnings PDFs, internal metric series, and board memos into a per-quarter HydraDB collection
- Answer trend questions like “How did gross margin change across the last four quarters?” using
recency_biasandmode: "thinking" - Compare point-in-time statements across quarters (“What did the CFO say about guidance in Q2 vs Q4?”)
- Build a financial analyst interface that cites the exact source document and timestamp for every answer
Why Naive RAG Fails on Financial Data
The structural problem is temporal ambiguity. A Q2 2023 earnings call and a Q4 2023 earnings call are nearly identical in vocabulary, format, and topic distribution. Both discuss revenue, margins, guidance, and macro headwinds. Their embeddings sit close together in vector space. When you ask “how did guidance change between Q2 and Q4?”, a cosine similarity search returns whichever call scores slightly higher: not both, not in order, not with any awareness that temporal comparison is what the question requires. HydraDB fixes this through three architectural properties:- Timestamp-aware indexing: each document you ingest can carry an ISO 8601 event date (
event_timeinadditional_metadata) that HydraDB reads as a first-class attribute alongside the vector.recency_biasuses this to weight results by recency or spread results across time depending on what the query needs. - Temporal Knowledge Graph: entities (companies, executives, metrics, products) are stored as nodes. Each mention of a metric across different documents creates a time-ordered edge sequence on that entity, effectively a versioned history. Querying “revenue trend” traverses these edges in temporal order rather than returning a flat ranked list of chunks.
- Multi-query expansion: in
mode: "thinking", HydraDB expands the query into semantically diverse reformulations, then executes all of them in parallel. “How did guidance change between Q2 and Q4?” becomes “Q2 guidance outlook”, “Q4 forward guidance revised”, “guidance comparison quarterly”, each targeting a different point on the timeline.
Architecture Overview
Key design decisions:- One database for the entire financial data corpus. Collections namespace by company or analyst team.
recency_bias: 0.7for point-in-time questions (target the quarter but still gather adjacent context).recency_bias: 0.3for trend and historical comparison (spread evenly across time).mode: "thinking"for all analytical queries: multi-query reranking is essential for financial reasoning.graph_context: trueon/querygives youquery_paths: the entity relationship chains that show how a metric’s value evolved across documents.
Step 1: Create Database & Environment
One database for the whole financial corpus. Collections isolate by company, fund, or analyst team. Declare the fields you filter on (ticker, doc_type, fiscal_year, fiscal_quarter) in the metadata schema now: top-level metadata_filters match schema fields, and schema field names cannot be renamed later.
SDK used for all API calls in this guide. Install:pip install hydradb-sdk. Import asfrom hydra_db import HydraDB(the import name differs from the package name).
Step 2: Upload Financial Documents
2.1 Earnings Call Transcripts & SEC Filings (PDF)
Earnings PDFs are the primary context. Tag each with structured metadata (ticker, doc_type, fiscal_year, fiscal_quarter) so you can filter search to a specific company or exact fiscal period before semantic search even runs. For an uploaded file, per-file metadata goes in the document_metadata form field: a JSON array with one item per file, in the same order as documents.
metadata: the fields you filter on (ticker,doc_type,fiscal_year,fiscal_quarter), declared in the Step 1 schema. In a database with a schema, an undeclaredmetadatakey is rejected.additional_metadata: labels you read back (period_label,period_end_date) andevent_time. Recency ranking reads the event time fromadditional_metadata(event_time,timestamp, ordate), never the upload time, so set it to the period end.- Title:
document_metadatatakes notitle. The source title comes from the filename, so name files clearly (for exampleNVDA-Q2-2023-earnings.pdf).
Theevent_timefield is load-bearing. HydraDB usesadditional_metadata.event_timeto weight results whenrecency_biasis set. If you leave it blank or use the upload date instead of the reporting period end date, every Q2 and Q4 call will look equally “recent” and temporal queries will fail. Always use the fiscal period end date.
2.2 Internal Metrics (CSV / JSON)
Internal financial metrics (revenue, ARR, churn, CAC, LTV, burn rate) are typically exported from a data warehouse or BI tool as CSV or JSON. Convert them to structured text chunks with one row per metric-per-period before uploading. This gives HydraDB the granularity to answer “what was CAC in Q2 2023?” precisely.2.3 Board Memos & Investor Letters
Board memos contain the strategic narrative behind the numbers: the reasoning that doesn’t appear in the income statement. Upload them alongside the earnings PDFs. HydraDB’s context graph automatically links board memo references to related earnings call chunks.2.4 Verify Indexing Before Going Live
Always verify all uploaded documents are indexed before running any queries. Unverified documents return silently empty results, a hard bug to diagnose in production.Pacing reminder. The batch helper in 2.1 pauses 1 second after every 20 uploads. Call/context/statusbefore any production query. Anerroredstatus witherror_code: "FILE_NOT_FOUND"usually means thecollectionyou asked about is not the one you ingested into.
Step 3: Store Analyst Memory
Per-analyst memory records the analyst’s focus area, preferred companies, and communication style, which your app passes to the LLM to personalize answers. A macro fund PM cares about different metrics than a sector-specialist equity analyst.Setinfer: truefor analyst profiles (the default isfalse). HydraDB extracts signals likeuser COVERS ticker:ACME,user PREFERS format:quantitative,user FOCUS unit_economics, and builds graph connections automatically. The profile lives in the analyst’s own collection, so knowledge queries on other collections do not read it. Steps 4.2 and 6 fetch it with a separatetype: "memory"query onanalyst-{user_id}and pass it to the LLM.
Step 4: Temporal Search Queries
This is the core of the financial analyst use case. Four distinct query patterns, each with differentrecency_bias and retrieval configuration.
4.1 Point-in-Time: “What happened in Q2?”
Use highrecency_bias to surface the most relevant recent documents. For point-in-time questions, also use metadata_filters to scope to the exact quarter. The transcripts are knowledge and the metrics are memories in different collections, so the query uses type="all" with collections to read both.
4.2 Trend Analysis: “How did X change across quarters?”
For trend questions, remove the quarter filter and lowerrecency_bias so HydraDB spreads results across the full timeline. This is the key pattern that naive RAG gets wrong.
Whyrecency_bias: 0.3for trend queries? Withrecency_bias: 0.7, HydraDB weights recent quarters heavily: you get Q4 results dominating, and Q1/Q2 are underrepresented. For trend analysis you need all four quarters equally weighted. Settingrecency_bias: 0.3spreads retrieval across the timeline without penalising recent data entirely.
4.3 Cross-Source Synthesis: “Do the numbers match the narrative?”
This is Priya’s use case: reconciling what the CEO said on the earnings call against internal metrics and the board memo for the same quarter. HydraDB’s context graph automatically links the three sources by entity (the company, the metric, the period).4.4 Guidance Tracking: “How has management’s tone on X changed?”
Track how management’s language around a specific topic (e.g. guidance, macro risk, hiring) has shifted across quarters. Usesrecency_bias: 0.2 and full timeline retrieval, with explicit chronological sorting. The code sets graph_context: false, but that flag only takes effect in mode: "fast"; with mode: "thinking", as here, the graph slice is always included.
Step 5: Analyst Search Interface
For analyst chat interfaces, usePOST /query to retrieve chunks, sources, and graph context. Generate the final answer in your application layer with your LLM provider so citations, formatting, and conversation memory stay under your control.
Step 6: Automated Quarterly Briefing Agent
Run this agent after each earnings release. It assembles a full briefing (performance summary, trend table, guidance revision, and narrative shift) and saves it back to HydraDB as a memory for the analyst’s next session.Step 7: Multi-Analyst Slack Interface (Optional)
Expose the analyst to your team via a Slack slash command. The handler maps each Slack user to an analyst, pulls the ticker and quarter out of the question, and picksrecency_bias from the question’s wording.
Complete API Reference
All endpoints used in this cookbook. Base URL:https://api.hydradb.comHeader:
Authorization: Bearer YOUR_API_KEY
Create Database
GET /databases/status?database=financial-analyst until data.infra.ready_for_ingestion is true.
Upload Financial Document (SDK)
Upload PDF via cURL
Verify Indexing
Store Metrics / Analyst Memory
Point-in-Time Search
Trend Search (raw chunks + graph paths)
Historical Comparison (very low recency_bias)
Search Analyst Preferences
Response Shape: /query
This object is the data field of the response envelope (success, data, error, meta), which the SDK exposes as result.data.
Recency Bias Quick Reference
Benchmarks
Tested across 3 company corpora (4 quarters of earnings transcripts + internal metrics + board memos each). Compared against naive vector RAG baseline and a manual analyst workflow.
On the 18% trend accuracy for naive RAG. This is a structural limitation. Embedding a Q2 earnings call and a Q4 earnings call produces very similar vectors: they are the same format, vocabulary, and topic distribution. Vector search returns whichever ranks slightly higher, ignoring the other entirely. HydraDB’s timestamp-aware retrieval, recency_bias, and multi-query expansion solve this without any prompt engineering.
Benchmark methodology. Figures are based on internal HydraDB testing. For the formal benchmark paper and methodology, see research.hydradb.com/hydradb.pdf. Results will vary by corpus size, document quality, and query distribution.
Common Pitfalls & Fixes
Next Steps
- Expand your corpus: add 10-K and 10-Q filings as
doc_type: "10K"/"10Q", withevent_timeset to the filing period end. - Add a second company: create a new collection
earnings-{TICKER2}and compare two companies directly using cross-ticker queries. - Schedule automated ingestion: run the ingestion pipeline on a cron job triggered by each earnings release date.
- Wire up a Slack briefing: schedule
generate_quarterly_briefingto post to a#earnings-briefingschannel automatically after each filing. - Add a web scraper: ingest sell-side analyst notes or financial news articles alongside the official filings for a richer context graph.
