Unstructured

Research Team Data Published Compare

Summary

GenAI data platform that transforms complex, unstructured and multimodal data into clean, structured, AI-ready inputs for RAG, search, analytics, and LLM applications.

Description

Processes 64+ file types with connectors, partitioning, VLM parsing, chunking, enrichment, embedding, and workflow/API endpoints to prepare enterprise data for RAG and GenAI pipelines.

Positioning

Data preparation layer for GenAI and RAG

Key facts

HQ location
San Francisco, CA, USA
Founded
2020
Employee range
51-200
Funding stage
Series B
Company type
Private (Private / open-source commercial)
Pricing model
Licensing Open Source Subscription Usage Based (Open-source plus API/SaaS usage and enterprise plans)
Last updated

Financials

Revenue estimate
$7.7M ARR estimate (Latka, Jan 2025); company revenue not officially disclosed
Valuation estimate
~$230M reported/estimated Series B valuation
Investments
$65M total; $40M Series B announced Mar 2024

Relationships

Target customers
Enterprise AI teams, developers, data teams, companies building RAG and knowledge applications
Key competitors
LlamaIndex, LangChain, Upstage, Haystack/deepset, Pinecone, Weaviate, Databricks
Known customers
Not publicly disclosed

Classification (raw research text)

Core focus
AI data ingestion and preprocessing
Core industry
AI Data Infrastructure
Core category
RAG data preparation platform

Shown verbatim from the research spreadsheet — deriving structured industry tags from this text is a future phase.

Segments, Industries & Certifications

Segments

AI Workflows, AI Developer Tools, Knowledge & RAG, Document AI, AI Governance & Risk, Chatbots, Analytics & BI