Summary
GenAI data platform that transforms complex, unstructured and multimodal data into clean, structured, AI-ready inputs for RAG, search, analytics, and LLM applications.
Description
Processes 64+ file types with connectors, partitioning, VLM parsing, chunking, enrichment, embedding, and workflow/API endpoints to prepare enterprise data for RAG and GenAI pipelines.
Positioning
Data preparation layer for GenAI and RAG
Key facts
- HQ location
- San Francisco, CA, USA
- Founded
- 2020
- Employee range
- 51-200
- Funding stage
- Series B
- Company type
- Private (Private / open-source commercial)
- Pricing model
- Licensing Open Source Subscription Usage Based (Open-source plus API/SaaS usage and enterprise plans)
- Last updated
Financials
- Revenue estimate
- $7.7M ARR estimate (Latka, Jan 2025); company revenue not officially disclosed
- Valuation estimate
- ~$230M reported/estimated Series B valuation
- Investments
- $65M total; $40M Series B announced Mar 2024
Relationships
- Target customers
- Enterprise AI teams, developers, data teams, companies building RAG and knowledge applications
- Key competitors
- LlamaIndex, LangChain, Upstage, Haystack/deepset, Pinecone, Weaviate, Databricks
- Known customers
- Not publicly disclosed
Classification (raw research text)
- Core focus
- AI data ingestion and preprocessing
- Core industry
- AI Data Infrastructure
- Core category
- RAG data preparation platform
Shown verbatim from the research spreadsheet — deriving structured industry tags from this text is a future phase.
Segments, Industries & Certifications
- Segments