diff --git a/README.md b/README.md index 6bc7fd6..63c21e4 100644 --- a/README.md +++ b/README.md @@ -16,6 +16,67 @@ No required OCR stack. No document-processing service lock-in. ## Quick Start +Install doc7, start a local vision model in LM Studio or Ollama, then convert a +document: + +```bash +# macOS or Linux +curl -fsSL https://raw.githubusercontent.com/magicrew/doc7/main/scripts/install.sh | bash + +# Windows PowerShell +irm https://raw.githubusercontent.com/magicrew/doc7/main/scripts/install.ps1 | iex + +# Convert a document +doc7 report.pdf +``` + +The first run discovers the local model endpoint and saves the selected model +on this machine. For a remote endpoint, configure it with `doc7 setup`. + +## Understand Complex Pages + +[![doc7 turns Attention Is All You Need into AI-ready Markdown](./examples/attention-is-all-you-need/showcase.webp)](./examples/attention-is-all-you-need/input.webp) + +One raster-only page from *Attention Is All You Need*. No text layer. doc7 +recovers the paper identity, Figure 2, the displayed equation, the scaling +rationale, the technical footnote, and the ordered relationships inside both +attention diagrams as searchable Markdown. + +The same pipeline processes complete papers and multi-page reports, then +rebuilds the ordered pages into one document. + +`doc7` reads the whole page instead of stopping at character extraction. Bring +any OpenAI-compatible multimodal model, including a private or local deployment. +There is no required OCR stack and no per-page document-parser fee from doc7. + +## Measured on Real Inputs + +[![doc7 Visual Understanding Benchmark](./assets/readme/benchmark/benchmark.webp)](./benchmarks/visual-report/README.md) + +Two raster-only PDFs. Fifteen machine-checkable visual facts. MarkItDown OCR +and doc7 used the same `qwen3.5-9b` model through the same local +OpenAI-compatible endpoint. Docling used its standard local pipeline. + +In this run, doc7 recovered **15/15** checked facts, compared with 9/15 for +MarkItDown with its OCR plugin and 3/15 for Docling's standard pipeline. + +## One Pipeline for Every Format + +[![doc7 processes every document through one visual-understanding pipeline](./assets/readme/formats/formats.webp)](#supported-inputs) + +Different containers enter the same page-understanding pipeline. Text, tables, +formulas, charts, diagram relationships, image meaning, and visible UI state +leave as one searchable Markdown document. + +## Built Around the CLI + +[![doc7 command-line interface on macOS, Linux, and Windows](./assets/readme/cli/cli.webp)](#quick-start) + +The same binary provides the interactive CLI, batch processing, model checks, +MCP, the Go SDK, and the asynchronous HTTP service. + +## Detailed Setup and Usage + **Download:** [macOS, Linux, and Windows CLI archives](https://github.com/magicrew/doc7/releases) Install the latest release without administrator privileges. @@ -200,27 +261,7 @@ operate a model. doc7 is designed for the opposite case: reuse a local or private model, eliminate a recurring document-parser bill, and turn long-term document conversion into infrastructure you own. -## High-Complexity Papers, Precise Markdown - -[![doc7 turns Attention Is All You Need into AI-ready Markdown](./examples/attention-is-all-you-need/showcase.webp)](./examples/attention-is-all-you-need/input.webp) - -One raster-only page from *Attention Is All You Need*. No text layer. doc7 -recovers the paper identity, Figure 2, the displayed equation, the scaling -rationale, the technical footnote, and the ordered relationships inside both -attention diagrams as searchable Markdown. - -The same pipeline processes complete papers and multi-page reports, then -rebuilds the ordered pages into one document. - -`doc7` reads the whole page instead of stopping at character extraction. Bring any OpenAI-compatible multimodal model, including a private or local deployment. There is no required OCR stack and no per-page document-parser fee from doc7. - -## Open Benchmark - -[![doc7 Visual Understanding Benchmark](./assets/readme/benchmark/benchmark.webp)](./benchmarks/visual-report/README.md) - -Two raster-only PDFs. Fifteen machine-checkable visual facts. MarkItDown OCR -and doc7 used the same `qwen3.5-9b` model through the same local -OpenAI-compatible endpoint. Docling used its standard local pipeline. +## Open Benchmark Details | System | Attention paper | Visual report | Combined | Raw Markdown | | --- | ---: | ---: | ---: | ---: | @@ -253,14 +294,6 @@ page 4 of *Attention Is All You Need* and is excluded from doc7's MIT license. For a pinned large-scale evaluation, use the [olmOCR-Bench adapter](./benchmarks/olmocr/README.md). It supports the upstream 1,403-PDF / 7,010-fact suite without redistributing its third-party documents; doc7 does not claim a full-suite ranking until a complete pinned run is published. -## One Pipeline, Every Document - -[![doc7 processes every document through one visual-understanding pipeline](./assets/readme/formats/formats.webp)](#supported-inputs) - -Different containers enter the same page-understanding pipeline. Text, -tables, formulas, charts, diagram relationships, image meaning, and visible UI -state leave as one searchable Markdown document. - ## A Different Architecture | Primary approach | Representative projects | What happens | What you operate | @@ -272,13 +305,6 @@ state leave as one searchable Markdown document. The model remains your choice. `doc7` does not prescribe a model size or claim quality that has not been measured. With a private open model, there is no required external document-processing service and no doc7 usage quota. -## The CLI Is The Product - -[![doc7 command-line interface on macOS, Linux, and Windows](./assets/readme/cli/cli.webp)](#quick-start) - -The same binary provides the interactive CLI, batch processing, model checks, -MCP, the Go SDK, and the asynchronous HTTP service. - ## Model, Dependency, and Recovery Workflow Discover the model IDs exposed by an OpenAI-compatible vision endpoint: diff --git a/README.zh-CN.md b/README.zh-CN.md index d539176..9f7b9ae 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -16,6 +16,60 @@ doc7 通过你自己的 OpenAI 兼容多模态模型,把 PDF、Office、扫描 ## 快速开始 +安装 doc7,启动 LM Studio 或 Ollama 中的本地视觉模型,然后转换文档: + +```bash +# macOS 或 Linux +curl -fsSL https://raw.githubusercontent.com/magicrew/doc7/main/scripts/install.sh | bash + +# Windows PowerShell +irm https://raw.githubusercontent.com/magicrew/doc7/main/scripts/install.ps1 | iex + +# 转换文档 +doc7 报告.pdf +``` + +首次运行会自动发现本地模型接口,并把选择保存到当前机器。需要使用远程接口 +时,运行 `doc7 setup` 配置即可。 + +## 复杂页面,也能完整理解 + +[![doc7 将 Attention Is All You Need 转换为 AI 可用的 Markdown](./examples/attention-is-all-you-need/showcase.zh-CN.webp)](./examples/attention-is-all-you-need/input.webp) + +《Attention Is All You Need》中的一页纯图片 PDF,没有文本层。doc7 把论文身份、 +Figure 2、展示公式、缩放原因、技术脚注,以及两张注意力图内部的顺序和并行关系, +完整转换为可检索的 Markdown。 + +同一套流程也可以处理完整论文和多页报告,再按原始页序重建为一份文档。 + +doc7 读取整页信息,而不止字符。你可以接入任何兼容 OpenAI 接口的多模态模型, +包括本地模型和私有化部署。doc7 不要求单独搭建 OCR 技术栈,也不收取按页处理费用。 + +## 用真实输入衡量效果 + +[![doc7 视觉理解 Benchmark](./assets/readme/benchmark/benchmark.zh-CN.webp)](./benchmarks/visual-report/README.md) + +两份纯图片 PDF,15 项可机器校验的视觉事实。MarkItDown OCR 与 doc7 通过同一个 +本地 OpenAI 兼容接口使用同一个 `qwen3.5-9b` 模型,Docling 使用标准本地流水线。 + +这次运行中,doc7 恢复了 **15/15** 项事实;MarkItDown 加 OCR 为 9/15,Docling +标准流水线为 3/15。 + +## 一套流程处理所有格式 + +[![doc7 使用一套视觉理解流程处理一切文档](./assets/readme/formats/formats.zh-CN.webp)](#支持的输入) + +不同格式进入同一套页面理解流程。正文、表格、公式、图表、图示关系、图像含义和 +可见界面状态,最终组成一份可检索的 Markdown 文档。 + +## 从命令行直接开始 + +[![doc7 在 macOS、Linux 和 Windows 上的命令行界面](./assets/readme/cli/cli.zh-CN.webp)](#快速开始) + +同一个二进制文件提供交互式 CLI、批量处理、模型检查、MCP、Go SDK 和异步 HTTP 服务。 + +## 详细配置与使用 + **直接下载:** [macOS、Linux 和 Windows CLI 压缩包](https://github.com/magicrew/doc7/releases) 不需要管理员权限,直接安装最新发行版。 @@ -176,25 +230,7 @@ API 通常按页、图片或功能调用计费;本地 doc7 的主要成本是 选择:复用本地或私有模型,取消持续累加的文档解析账单,把长期文档转换能力 变成自己拥有的基础设施。 -## 高复杂度论文,也能精准转换 - -[![doc7 将 Attention Is All You Need 转换成 AI 可用的 Markdown](./examples/attention-is-all-you-need/showcase.zh-CN.webp)](./examples/attention-is-all-you-need/input.webp) - -《Attention Is All You Need》中的一页纯图片 PDF,没有文本层。doc7 -把论文身份、Figure 2、展示公式、缩放原因、技术脚注,以及两张注意力图 -内部的顺序和并行关系,完整转换为可检索的 Markdown。 - -同一套流程也可以处理完整论文和多页报告,再按原始页序重建为一份文档。 - -doc7 读取整页信息,而不止字符。你可以接入任何兼容 OpenAI 接口的多模态模型,包括本地模型和私有化部署。doc7 不要求单独搭建 OCR 技术栈,也不收取按页处理费用。 - -## 公开 Benchmark - -[![doc7 视觉理解 Benchmark](./assets/readme/benchmark/benchmark.zh-CN.webp)](./benchmarks/visual-report/README.md) - -两份纯图片 PDF,15 项可机器校验的视觉事实。MarkItDown OCR 与 doc7 -通过同一个本地 OpenAI 兼容接口使用同一个 `qwen3.5-9b` 模型,Docling -使用标准本地流水线。 +## 公开 Benchmark 详情 | 系统 | Attention 论文 | 视觉报告 | 合计 | 原始 Markdown | | --- | ---: | ---: | ---: | ---: | @@ -223,13 +259,6 @@ doc7 的 MIT 许可证范围。[查看来源和授权记录](./examples/attentio 需要大规模评估时,使用 [olmOCR-Bench 适配器](./benchmarks/olmocr/README.md)。它支持上游固定版本的 1,403 份 PDF、7,010 个机器可判定事实,不把第三方文档打包进 doc7;在完整、固定版本的运行完成前,doc7 不发布全量排名结论。 -## 一套流程,处理一切文档 - -[![doc7 使用一套视觉理解流程处理一切文档](./assets/readme/formats/formats.zh-CN.webp)](#支持的输入) - -不同格式进入同一套页面理解流程。正文、表格、公式、图表、图示关系、图像 -含义和可见界面状态,最终组成一份可检索的 Markdown 文档。 - ## doc7 选择了另一条路线 | 主要路线 | 代表项目 | 处理方式 | 需要维护的东西 | @@ -241,13 +270,6 @@ doc7 的 MIT 许可证范围。[查看来源和授权记录](./examples/attentio 模型由用户决定。doc7 不预设模型规模,也不会宣传没有实测过的模型效果。接入私有化开源模型后,不再需要购买专门的文档解析服务,也没有 doc7 处理额度。 -## 命令行就是产品入口 - -[![doc7 在 macOS、Linux 和 Windows 上的命令行界面](./assets/readme/cli/cli.zh-CN.webp)](#快速开始) - -同一个二进制文件提供交互式 CLI、批量处理、模型检查、MCP、Go SDK 和异步 -HTTP 服务。 - ## 模型、依赖与失败恢复 先查看兼容 OpenAI 的视觉模型接口实际提供的模型 ID: