Data & Ingestion

Getting real-world content into a form an agent can use — crawling, scraping, parsing PDFs and normalising documents.

12 projects

FirecrawlData

Renders JavaScript, follows links and strips the navigation and boilerplate that otherwise poison a retrieval corpus. Solves the unglamorous half of RAG properly. AGPL-licensed, so check the terms if you plan to expose it as a service.

OfficialAGPL-3.0174k+3.2k
ragflowRAG & Retrieval

RAGFlow is a server you run rather than a library you import: documents go into a dataset, it splits them into chunks, and you can read every chunk on screen and hand-correct the ones that came out wrong before any of it reaches a model, with chat answers citing the chunk they came from. Around that sit an agent canvas and an ingestion pipeline you can rebuild yourself, so retrieval, tools and MCP servers can be wired together without writing code. The optical character recognition, table-structure and layout models that make sense of scans run on PDFs and images only — Word, Excel and PowerPoint files are read structurally, and the docs tell you to convert a DOCX to PDF if you want the visual parser on it. The price of all this is weight: a Docker Compose stack with Elasticsearch, MySQL, MinIO and Redis behind it, 4 cores and 16 GB of RAM as the stated minimum, and parsing that the project's own FAQ concedes is slower than LangChain's. If your documents are already clean text and you want retrieval inside your own Python application, LlamaIndex or Haystack will fit better.

OfficialApache-2.089.5k+546
PaddleOCRData

Retrieval over your own documents usually fails long before the model does, at the point where a scanned table or a two-column PDF becomes an unreadable wall of text. PaddleOCR is the layer that stops that. It recognises text in more than a hundred languages, and its document pipeline preserves what the layout meant — headings stay headings, a table comes out as a table with its cell positions, formulas and seals are handled — emitting Markdown or JSON. The models are deliberately small enough to run on a CPU or at the edge, which is what makes processing a large archive affordable. The cost is the framework underneath: this is a PaddlePaddle project, and installing it means installing that.

Apache-2.088.3k
Crawl4AIData

A permissively licensed, self-hosted alternative for the same job as the hosted crawlers, with extraction strategies you can define per site. Being a library rather than a service means you own the proxy handling and rate limiting yourself.

OfficialApache-2.079.7k+763
ScraplingData

You pick a fetcher per job — plain HTTP with TLS fingerprint impersonation when speed matters, a Playwright-driven browser for JavaScript pages, or a hardened stealth browser that patches headless tells and solves Cloudflare Turnstile — then select elements six ways, from CSS and XPath to text and structural similarity. Adaptive mode fingerprints matched elements so the same call finds them again after a site redesign, and the spider added in 0.4 crawls with deduplication, per-domain limits, pause-and-resume and JSON or CSV export. An MCP server lets a coding agent scrape conversationally, trimming pages by selector before they reach the model. It is a structured-extraction library for one machine: whole-site-to-markdown corpus building is Crawl4AI's territory, and the Turnstile-solving stealth sits in open tension with the README's own education-and-research disclaimer.

BSD-3-Clause76.9k+1.3k
DoclingData

The difference from a text extractor is that structure survives: page layout, reading order, table cells, code blocks and formulas come through as parts of a single document representation rather than as a flattened string, which is what decides whether a table in a report is still a table by the time an agent reads it. Everything can run locally, including in air-gapped environments. The cost of that fidelity is compute — machine-learning models run over each page, so this is a pipeline to budget for, not a fast converter to call inline.

OfficialMIT65.7k+335
Scientific Agent SkillsTools & Integrations

An agent can already call any Python library, but it will use a specialist bioinformatics package the way someone reads the docs for the first time — plausibly, and wrong in the ways that matter. This supplies the missing half: over 160 skills that carry the curated usage and worked examples for specific scientific tools, plus access to a hundred or so public databases, across genomics, chemistry, drug discovery, proteomics, clinical research, imaging, physics and geospatial work. It installs as one plugin or skill by skill. Read the scope notes before relying on it, though — the clinical and healthcare skills are written for research and retrospective validation, and say plainly that they are not for patient-specific decisions.

Agent skillMIT36.8k
SearXNGData

It keeps no index of its own; a query is fanned out to as many as 269 search services and the results are merged, which is how an agent gets web search without a commercial search API behind it. A JSON endpoint makes it usable from code rather than only from the browser. Two consequences follow from having no index: result quality and availability are inherited from the upstream engines, and when one of them starts blocking your instance that is your problem to solve.

OfficialAGPL-3.036.2k+334
GraphRAGRAG & Retrieval

Ordinary retrieval finds the chunks most similar to a question, which fails when the answer is not written down in any one chunk — what the main themes are, how two people are connected. GraphRAG instead extracts entities, relationships and claims from every passage, clusters the resulting graph into communities and summarises each one, then answers from those summaries. The price is paid at indexing time: a model runs over every text unit and again over every community, so building the index scales with the size of the corpus, not with how often you query it.

OfficialMIT35.7k+106
GraphitiMemory & Context

Most memory layers overwrite: a new fact replaces the old one and last year's answer becomes unrecoverable. Graphiti instead marks the old relationship as no longer holding, with the dates attached, so both what is true now and what was true then remain queryable — and it separates when something happened from when the system learned about it, which is what keeps late-arriving information from rewriting history. Retrieval combines vector similarity, full-text search and graph traversal. The costs are concrete: a graph database has to run alongside your stack, and every ingest calls a model to extract entities and relationships.

OfficialApache-2.030.4k+217
skillsTools & Integrations

Over a hundred skill directories of Google's own written guidance, installed selectively with `npx skills add google/skills`. The bulk sits under `skills/cloud` — GKE, BigQuery, Agent Platform, the six Well-Architected pillars — while `skills/ads` and `skills/analytics` cover the Google Ads API, the Mobile Ads and IMA SDKs and the two Google Analytics APIs. Depth is uneven: `skills/cloud/gke-*` accounts for twenty-nine directories on its own, Cloud Run and Firebase get one each, and the README labels the repository as under active development. A stack on neither Google Cloud nor Google's ads and analytics products gets nothing here, and app-side coverage stops at the Mobile Ads SDK — Android, Flutter, Dart, Genkit and ADK skills are separate repositories the README only links to.

Agent skillOfficialApache-2.018.9k+332
Skill SeekersTools & Integrations

An agent working against an unfamiliar library reads the same documentation site over and over, a page at a time, paying for it on every run. Skill Seekers converts that source once: point it at a documentation site, a repository, a PDF, a wiki or a video and it produces a structured knowledge asset, which it can then export as an agent skill or into a retrieval pipeline. Eighteen kinds of source go in and it can package for a couple of dozen destinations, so the same extraction serves a coding assistant and a search index. It also checks for conflicts, which is the failure mode when two versions of the same documentation end up loaded at once.

MIT14.9k