S

Sycamore

by Aryn (acquired by Glean, March 2026)

Data & AnalyticsEnterprise Search & KnowledgeDeveloper Tools

Open-source document ETL and analytics engine for unstructured data

Free·Added Jul 10, 2026·Updated Sep 7, 2026
Share:
THE DAILY BRIEF
Sycamore

by Aryn (acquired by Glean, March 2026)

Data & AnalyticsEnterprise Search & KnowledgeDeveloper Tools

Open-source document ETL and analytics engine for unstructured data

Free

Sycamore is an Apache-2.0 document processing engine that partitions, enriches, chunks, embeds and loads complex unstructured documents - PDFs, presentations, manuals and transcripts with embedded tables and figures - into vector databases. It is built for data engineers assembling RAG and analytics pipelines who find that naive text splitters destroy document structure.

At a Glance

Category
Data & Analytics
Pricing
Free
Target Market
Data Scientists, Enterprise Developers, CTOs, CIOs
Deployment
Open-source, Self-hosted
Founded
2023

Key Features

  • ✓DocSet dataflow abstraction
  • ✓Structure-aware partitioning
  • ✓LLM-powered transforms
  • ✓Seven vector and search connectors
  • ✓Ray-backed scaling
  • ✓Local or hosted partitioning
  • ✓Configurable chunking strategies

Capabilities

✗text generation
✗image generation
✗video generation
✗code generation
✓workflow automation
✓api access
✗audio generation
✗fine tuning
✗agent orchestration

Use Cases

  • •Enterprise RAG over technical PDFs
  • •Analytics over report collections
  • •Weaviate ingestion pipeline
  • •OpenSearch RAG on AWS
  • •Bulk metadata extraction

Ideal For

Best For

  • ✓Structure-aware chunking of PDFs containing tables, figures and multi-column layouts
  • ✓Loading enriched document collections into OpenSearch, Weaviate, Pinecone, Qdrant, Elasticsearch, DuckDB or Neo4j
  • ✓LLM-powered entity and schema extraction as an explicit, testable pipeline stage rather than an ad-hoc prompt
  • ✓Scaling the same Python dataflow from a laptop to a Ray cluster without rewriting it
  • ✓Teams that need an auditable, fully open-source document pipeline for compliance or vendor-independence reasons

Not Ideal For

  • ✗Teams wanting a managed SaaS: Aryn was acquired by Glean and said its hosted cloud APIs would only run until 15 April 2026, so the hosted DocParse path is no longer a safe long-term dependency
  • ✗Non-engineers - Sycamore is a Python library with a DocSet API and a Ray runtime, not a point-and-click ingestion UI
  • ✗Projects handling only plain text documents, where a simple text splitter costs far less to operate than a Ray-based dataflow engine

Market Analysis

Open-sourceDeveloper-firstEnterprise data engineering

Pros

  • ✓Apache-2.0 with an explicit post-acquisition statement that it 'will remain so', so adopters are not exposed to a licence change
  • ✓Independent write-ups from AWS and Weaviate document working end-to-end pipelines rather than repeating vendor claims
  • ✓Seven vector and search connectors plus S3 and HTTP crawlers cover most enterprise RAG destinations without custom glue code
  • ✓The architecture was peer-reviewed and published at CIDR 2025, which is unusual for a document ingestion library

Cons

  • ✗The vendor no longer exists independently - Aryn was acquired by Glean in March 2026 and its hosted cloud service was slated to stop on 15 April 2026
  • ✗The community is small: 607 GitHub stars, 72 forks and 57 open issues at the current repository snapshot
  • ✗The default partitioning path calls a hosted service; running entirely locally requires the separate local-inference extra and GPU capacity
  • ✗Weaviate's own guide notes that Sycamore's nested properties must be flattened into underscore-separated names because Weaviate only filters on top-level attributes
  • ✗No SOC 2, ISO 27001 or other certification attaches to the open-source library itself - compliance is entirely the adopter's responsibility

Pricing

Open source (Apache 2.0)

$0

  • ✓Full engine under Apache License v2.0
  • ✓All transforms and connectors included
  • ✓Local partitioning via the local-inference extra
  • ✓Self-hosted on your own hardware or Ray cluster

Sycamore itself is free under Apache License v2.0 and always has been; the only paid component was Aryn's hosted DocParse cloud service, which Aryn said would stop operating on 15 April 2026 following the Glean acquisition. Real cost today is the compute you run it on, GPU capacity if you use local partitioning, plus the LLM and embedding API calls your transforms make.

Security & Compliance

✗soc2
✗gdpr
✗hipaa
✗iso27001
✗sso
✗data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, weekly.

beri.net

Subscribe at beri.net/subscribe for weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Sycamore is an Apache-2.0 document processing engine that partitions, enriches, chunks, embeds and loads complex unstructured documents - PDFs, presentations, manuals and transcripts with embedded tables and figures - into vector databases. It is built for data engineers assembling RAG and analytics pipelines who find that naive text splitters destroy document structure.

Sycamore is an open-source, Apache-2.0 document processing engine for ETL, RAG and analytics over unstructured data, created by Aryn and developed in the aryn-ai/sycamore repository since July 2023. Its core abstraction is the DocSet, a distributed collection of documents where each document holds metadata plus an ordered list of elements such as tables, headings and images. The design is modelled on Apache Spark's dataflow style and executes on Ray, so the same script runs unchanged on a laptop or a multi-node cluster. A typical pipeline partitions a PDF into labelled elements, extracts entities and schema using LLM prompts, summarises images into searchable text, normalises values such as dates and locations, merges elements into chunks using strategies like GreedySectionMerger, computes embeddings and writes to a vector store. The transform catalogue includes Embed, Explode, Filter, FlatMap, Map, MapBatch, Materialize, Merge, Partition, ExtractEntity, Extract Schema, LLM Query, Sketch and Summarize. Connectors cover OpenSearch, Elasticsearch, Pinecone, Weaviate, Qdrant, DuckDB and Neo4j, alongside crawlers for Amazon S3 and HTTP. Partitioning defaults to Aryn DocParse, a GPU-backed service built on a DETR model trained on more than 80,000 enterprise documents that Aryn and AWS say delivers up to 6x more accurate chunking and 2x improved recall versus off-the-shelf systems; a local-inference install extra runs the partitioner on your own hardware instead. The architecture was published at CIDR 2025 as 'The Design of an LLM-powered Unstructured Analytics System', alongside a query planner called Luna. Aryn was acquired by Glean in March 2026 and said its cloud APIs and website would remain operational until 15 April 2026, while stating that Sycamore 'started as an 100% open source project with Apache License v2.0 and will remain so'. The repository is not archived and was still receiving commits in August 2026.

Ideal Buyer

Data engineering teams building RAG or document-analytics pipelines over complex PDFs, who need structure-aware chunking and want the pipeline expressed in code they own rather than inside a hosted black box.

Key Benefit

Structure-preserving document ETL - tables, figures and section hierarchy survive into the vector store - expressed in roughly twenty lines of Python and scalable from laptop to Ray cluster.

At a Glance

Category
Data & Analytics
Pricing
Free
Target Market
Data Scientists, Enterprise Developers, CTOs, CIOs
Deployment
Open-source, Self-hosted
Founded
2023

Key Features

  • ✓
    DocSet dataflow abstraction

    A Spark-style distributed collection of documents and ordered elements that composable transforms operate over at scale.

  • ✓
    Structure-aware partitioning

    Aryn DocParse segments and labels PDFs, runs OCR and extracts tables and images into structured JSON elements.

  • ✓
    LLM-powered transforms

    ExtractEntity, Extract Schema, LLM Query and Summarize apply model calls as first-class, composable pipeline stages.

  • ✓
    Seven vector and search connectors

    Writes to OpenSearch, Elasticsearch, Pinecone, Weaviate, Qdrant, DuckDB and Neo4j from the same DocSet writer.

  • ✓
    Ray-backed scaling

    The same script runs unchanged on a laptop or across a multi-node Ray cluster for large corpora.

  • ✓
    Local or hosted partitioning

    A local-inference install extra runs the partitioner on your own hardware instead of calling a cloud service.

  • ✓
    Configurable chunking strategies

    Mergers such as GreedySectionMerger recombine labelled elements into chunks that respect section boundaries rather than character counts.

Capabilities

✗text generation
✗image generation
✗video generation
✗code generation
✓workflow automation
✓api access
✗audio generation
✗fine tuning
✗agent orchestration

Use Cases

  • •
    Enterprise RAG over technical PDFs

    Ingest manuals and reports so that tables and figures stay retrievable instead of being flattened into unusable noise.

  • •
    Analytics over report collections

    The CIDR 2025 paper shows natural-language queries over NTSB accident reports achieving better accuracy than RAG on analytics questions.

  • •
    Weaviate ingestion pipeline

    Weaviate's own engineering blog walks through reading PDFs, enriching, embedding and loading in about twenty lines of code.

  • •
    OpenSearch RAG on AWS

    The AWS Big Data blog documents a DocParse plus Sycamore pipeline populating an Amazon OpenSearch Service index for retrieval.

  • •
    Bulk metadata extraction

    Pull titles, authors and custom schema fields from thousands of documents using LLM prompts as a single batch transform.

Ideal For

Best For

  • ✓Structure-aware chunking of PDFs containing tables, figures and multi-column layouts
  • ✓Loading enriched document collections into OpenSearch, Weaviate, Pinecone, Qdrant, Elasticsearch, DuckDB or Neo4j
  • ✓LLM-powered entity and schema extraction as an explicit, testable pipeline stage rather than an ad-hoc prompt
  • ✓Scaling the same Python dataflow from a laptop to a Ray cluster without rewriting it
  • ✓Teams that need an auditable, fully open-source document pipeline for compliance or vendor-independence reasons

Not Ideal For

  • ✗Teams wanting a managed SaaS: Aryn was acquired by Glean and said its hosted cloud APIs would only run until 15 April 2026, so the hosted DocParse path is no longer a safe long-term dependency
  • ✗Non-engineers - Sycamore is a Python library with a DocSet API and a Ray runtime, not a point-and-click ingestion UI
  • ✗Projects handling only plain text documents, where a simple text splitter costs far less to operate than a Ray-based dataflow engine

Integrations

✓SDK Available
SDK:Python

Deployment

✓On-Premise

Market Analysis

Open-sourceDeveloper-firstEnterprise data engineering

Pros

  • ✓Apache-2.0 with an explicit post-acquisition statement that it 'will remain so', so adopters are not exposed to a licence change
  • ✓Independent write-ups from AWS and Weaviate document working end-to-end pipelines rather than repeating vendor claims
  • ✓Seven vector and search connectors plus S3 and HTTP crawlers cover most enterprise RAG destinations without custom glue code
  • ✓The architecture was peer-reviewed and published at CIDR 2025, which is unusual for a document ingestion library

Cons

  • ✗The vendor no longer exists independently - Aryn was acquired by Glean in March 2026 and its hosted cloud service was slated to stop on 15 April 2026
  • ✗The community is small: 607 GitHub stars, 72 forks and 57 open issues at the current repository snapshot
  • ✗The default partitioning path calls a hosted service; running entirely locally requires the separate local-inference extra and GPU capacity
  • ✗Weaviate's own guide notes that Sycamore's nested properties must be flattened into underscore-separated names because Weaviate only filters on top-level attributes
  • ✗No SOC 2, ISO 27001 or other certification attaches to the open-source library itself - compliance is entirely the adopter's responsibility

Pricing

Open source (Apache 2.0)

$0

  • ✓Full engine under Apache License v2.0
  • ✓All transforms and connectors included
  • ✓Local partitioning via the local-inference extra
  • ✓Self-hosted on your own hardware or Ray cluster

Sycamore itself is free under Apache License v2.0 and always has been; the only paid component was Aryn's hosted DocParse cloud service, which Aryn said would stop operating on 15 April 2026 following the Glean acquisition. Real cost today is the compute you run it on, GPU capacity if you use local partitioning, plus the LLM and embedding API calls your transforms make.

Security & Compliance

✗soc2
✗gdpr
✗hipaa
✗iso27001
✗sso
✗data residency

Connect

Sources

This page was written from 6 sources, 5 on domains other than github.com.

  1. 1.github.com — sycamorevendor
  2. 2.sycamore.readthedocs.io — stable
  3. 3.aryn.ai — aryn transition
  4. 4.weaviate.io — sycamore and weaviate
  5. 5.aws.amazon.com — supercharge your rag applications with amazon opensearch ser
  6. 6.arxiv.org — 2409.00847
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe