OCR 与文档智能研究 Wiki

Tag: end-to-end

8 items with this tag.

  • 2026年7月25日

    dots.ocr: Multilingual Document Layout Parsing (1.2B vision encoder + 1.7B decoder, ~3B)

    • ocr
    • document-parsing
    • unified-vlm
    • layout
    • multilingual
    • end-to-end
    • frontend
    • backend
  • 2026年7月25日

    Ovis-OCR2: 0.8B End-to-End Document Parsing (OmniDocBench SOTA)

    • ocr
    • document-parsing
    • end-to-end
    • small-model
    • data-engine
    • synthetic-data
    • distillation
    • milestone
  • 2026年7月23日

    General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model (GOT)

    • ocr
    • ocr-20
    • unified
    • end-to-end
    • milestone
  • 2026年7月23日

    HunyuanOCR: Commercial-grade Lightweight 1B OCR VLM

    • ocr
    • vlm
    • small-model
    • unified
    • spotting
    • translation
    • end-to-end
  • 2026年7月23日

    Logics-Parsing: End-to-End LVLM with RL for Layout & Reading Order

    • ocr
    • document-parsing
    • reinforcement-learning
    • reading-order
    • layout
    • html-output
    • end-to-end
  • 2026年7月23日

    OCRVerse: Towards Holistic OCR in End-to-End

    • ocr
    • holistic-ocr
    • end-to-end
    • data-engineering
    • sft-rl
    • text-centric
    • vision-centric
  • 2026年7月23日

    POINTS-Reader: Distillation-Free Adaptation of VLMs for Document Conversion

    • ocr
    • document-parsing
    • distillation-free
    • synthetic-data
    • self-improvement
    • end-to-end
  • 2026年7月23日

    Qianfan-OCR: Unified End-to-End Document Intelligence (Layout-as-Thought)

    • ocr
    • document-intelligence
    • end-to-end
    • layout-as-thought
    • thinking
    • reading-order
    • omnidocbench

8 items with this tag.

  • 2026年7月25日

    dots.ocr: Multilingual Document Layout Parsing (1.2B vision encoder + 1.7B decoder, ~3B)

    • ocr
    • document-parsing
    • unified-vlm
    • layout
    • multilingual
    • end-to-end
    • frontend
    • backend
  • 2026年7月25日

    Ovis-OCR2: 0.8B End-to-End Document Parsing (OmniDocBench SOTA)

    • ocr
    • document-parsing
    • end-to-end
    • small-model
    • data-engine
    • synthetic-data
    • distillation
    • milestone
  • 2026年7月23日

    General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model (GOT)

    • ocr
    • ocr-20
    • unified
    • end-to-end
    • milestone
  • 2026年7月23日

    HunyuanOCR: Commercial-grade Lightweight 1B OCR VLM

    • ocr
    • vlm
    • small-model
    • unified
    • spotting
    • translation
    • end-to-end
  • 2026年7月23日

    Logics-Parsing: End-to-End LVLM with RL for Layout & Reading Order

    • ocr
    • document-parsing
    • reinforcement-learning
    • reading-order
    • layout
    • html-output
    • end-to-end
  • 2026年7月23日

    OCRVerse: Towards Holistic OCR in End-to-End

    • ocr
    • holistic-ocr
    • end-to-end
    • data-engineering
    • sft-rl
    • text-centric
    • vision-centric
  • 2026年7月23日

    POINTS-Reader: Distillation-Free Adaptation of VLMs for Document Conversion

    • ocr
    • document-parsing
    • distillation-free
    • synthetic-data
    • self-improvement
    • end-to-end
  • 2026年7月23日

    Qianfan-OCR: Unified End-to-End Document Intelligence (Layout-as-Thought)

    • ocr
    • document-intelligence
    • end-to-end
    • layout-as-thought
    • thinking
    • reading-order
    • omnidocbench

Created with Quartz © 2026

  • GitHub