OCR 与文档智能研究 Wiki

Tag: vlm

6 items with this tag.

  • 2026年8月01日

    视觉 Token 压缩(Visual Token Compression)

    • ocr
    • visual-token-compression
    • efficiency
    • kv-cache
    • vlm
    • back-ar-relevant
  • 2026年8月01日

    LayoutLite — Token 级隐式版面分析加速文档 OCR

    • ocr
    • document-parsing
    • visual-token-compression
    • layout-analysis
    • grpo
    • reinforcement-learning
    • vlm
    • kv-cache
    • front-detr-relevant
    • back-ar-relevant
    • plug-and-play
  • 2026年7月25日

    HPD-Parsing — 层次并行解码文档解析

    • ocr
    • document-parsing
    • decoding
    • parallel-decoding
    • mtp
    • speculative-decoding
    • kv-cache
    • vlm
    • baidu
    • front-detr-relevant
    • back-ar-relevant
  • 2026年7月23日

    HunyuanOCR: Commercial-grade Lightweight 1B OCR VLM

    • ocr
    • vlm
    • small-model
    • unified
    • spotting
    • translation
    • end-to-end
  • 2026年7月23日

    MinerU2.5: A Decoupled VLM for Efficient High-Resolution Document Parsing

    • ocr
    • document-parsing
    • vlm
    • decoupled
    • high-resolution
    • coarse-to-fine
    • frontend
    • backend
  • 2026年7月23日

    PaddleOCR-VL: 0.9B Ultra-Compact VLM for Document Parsing

    • ocr
    • document-parsing
    • vlm
    • small-model
    • layout
    • omnidocbench
    • frontend
    • backend

6 items with this tag.

  • 2026年8月01日

    视觉 Token 压缩(Visual Token Compression)

    • ocr
    • visual-token-compression
    • efficiency
    • kv-cache
    • vlm
    • back-ar-relevant
  • 2026年8月01日

    LayoutLite — Token 级隐式版面分析加速文档 OCR

    • ocr
    • document-parsing
    • visual-token-compression
    • layout-analysis
    • grpo
    • reinforcement-learning
    • vlm
    • kv-cache
    • front-detr-relevant
    • back-ar-relevant
    • plug-and-play
  • 2026年7月25日

    HPD-Parsing — 层次并行解码文档解析

    • ocr
    • document-parsing
    • decoding
    • parallel-decoding
    • mtp
    • speculative-decoding
    • kv-cache
    • vlm
    • baidu
    • front-detr-relevant
    • back-ar-relevant
  • 2026年7月23日

    HunyuanOCR: Commercial-grade Lightweight 1B OCR VLM

    • ocr
    • vlm
    • small-model
    • unified
    • spotting
    • translation
    • end-to-end
  • 2026年7月23日

    MinerU2.5: A Decoupled VLM for Efficient High-Resolution Document Parsing

    • ocr
    • document-parsing
    • vlm
    • decoupled
    • high-resolution
    • coarse-to-fine
    • frontend
    • backend
  • 2026年7月23日

    PaddleOCR-VL: 0.9B Ultra-Compact VLM for Document Parsing

    • ocr
    • document-parsing
    • vlm
    • small-model
    • layout
    • omnidocbench
    • frontend
    • backend

Created with Quartz © 2026

  • GitHub