82 items under this folder.2026年9月03日D-FINE — 用细粒度分布精修重定义 DETR 框回归object-detectiondetrreal-time-detectionbbox-regressiondistribution-refinementself-distillationlocalizationfront-detr-relevant2026年9月02日ArmorOCR — 困难视觉文本的定位、识别与 Spotting 联合强化ocrscene-texttext-spottinglocalizationadversarial-ocrself-distillationgrpobenchmarkfront-detr-relevantback-ar-relevant2026年9月02日SCVER — 解码状态条件的高分辨率视觉证据检索ocrdocument-parsingautoregressive-decodingvisual-retrievaldeformable-attentionhigh-resolutionefficiencyfront-detr-relevantback-ar-relevant2026年8月13日Muon: MomentUm Orthogonalized by Newton-Schulzoptimizertrainingorthogonalizationnewton-schulzspectralmilestoneinfra2026年8月06日Learning Transferable Visual Models From NL Supervision (CLIP)multimodalcvnlpmilestone2026年8月06日Deformable DETR: Deformable Transformers for End-to-End Object Detectiondetectiondetrdeformable-attentionmulti-scalefrontendmilestone2026年8月06日End-to-End Object Detection with Transformers (DETR)detectiondetrtransformerset-predictionmilestonefrontend2026年8月06日DINO: DETR with Improved DeNoising Anchor Boxesdetectiondetrdenoisingquery-selectionsotafrontendmilestone2026年8月06日Masked Autoencoders Are Scalable Vision Learners (MAE)sslcvtransformermilestone2026年8月06日An Image is Worth 16x16 Words (Vision Transformer, ViT)transformercvmilestone2026年8月01日LayoutLite — Token 级隐式版面分析加速文档 OCRocrdocument-parsingvisual-token-compressionlayout-analysisgrporeinforcement-learningvlmkv-cachefront-detr-relevantback-ar-relevantplug-and-play2026年8月01日MORE — 149 语种多语言文档解析基准ocrdocument-parsingbenchmarkmultilingualevaluationtedscdmtencentreading-orderfront-detr-relevant2026年7月29日BEiT: BERT Pre-Training of Image Transformersself-supervisedmimmasked-image-modelingvitpretraining2026年7月29日Extracting and Composing Robust Features with Denoising Autoencoders (DAE)self-superviseddenoisingrepresentation-learningfoundational2026年7月29日DiG: Reading and Writing — Discriminative and Generative Modeling for Self-Supervised Text Recognitionocrtext-recognitionself-supervisedmimcontrastive-learningmilestone2026年7月29日SeqCLR — Sequence-to-Sequence Contrastive Learning for Text Recognitionocrself-supervisedcontrastive-learningrepresentation-learninghtrctcattention2026年7月29日SimMIM: A Simple Framework for Masked Image Modelingself-supervisedmimmasked-image-modelingvitpretraining2026年7月28日OCR 系统 vs 多模态 LLM 的系统性评测(Caravani et al., 2027)ocrbenchmarkmllmevaluationlatencycostfine-tuningfca2026年7月25日ImageNet Classification with Deep CNNs (AlexNet)cnncvimagenetmilestone2026年7月25日DeepSeek-OCR: Contexts Optical Compressionocroptical-compressionvision-tokenmoemilestone2026年7月25日dots.ocr: Multilingual Document Layout Parsing (1.2B vision encoder + 1.7B decoder, ~3B)ocrdocument-parsingunified-vlmlayoutmultilingualend-to-endfrontendbackend2026年7月25日FireRed-OCR: Specializing General VLMs into OCR Modelsocrdocument-parsinggrpostructural-hallucinationdata-factorythree-stage-trainingmarkdown2026年7月25日GLM-OCR Technical Reportocrdocument-parsingcompactmulti-token-predictionmilestone2026年7月25日HPD-Parsing — 层次并行解码文档解析ocrdocument-parsingdecodingparallel-decodingmtpspeculative-decodingkv-cachevlmbaidufront-detr-relevantback-ar-relevant2026年7月25日HSD — 免训练层次投机解码加速文档解析ocrdocument-parsingdecodingspeculative-decodingparallel-decodingtraining-freekv-cacheback-ar-relevant2026年7月25日MonkeyOCRv2: A Visual-Text Foundation Model for Document AIocrdocument-aivisual-encoderfoundation-modelmilestone2026年7月25日Nougat: Neural Optical Understanding for Academic Documentsdocument-aiocrformulamilestone2026年7月25日Ovis-OCR2: 0.8B End-to-End Document Parsing (OmniDocBench SOTA)ocrdocument-parsingend-to-endsmall-modeldata-enginesynthetic-datadistillationmilestone2026年7月25日Parser-Oriented Structural Refinement — 稳定 DETR→解析器 界面ocrdocument-parsingdetrd-finelayout-interfaceretentionreading-ordernms-freefront-detr-relevantback-ar-relevant2026年7月25日RT-DocLayout — 实时端到端文档版面分析(含阅读顺序)ocrdocument-layout-analysisdladetrrt-detrdetectionsegmentationreading-orderbaidufront-detr-relevant2026年7月25日Unlimited OCR — R-SWA 恒定 KV 长文档解析ocrdocument-parsingkv-cachesliding-window-attentionlong-documentdecodingbaiduback-ar-relevant2026年7月23日Attention Is All You Need (Transformer)transformerattentionnlpfoundationmilestone2026年7月23日BERT: Pre-training of Deep Bidirectional Transformersnlptransformerpretrainingmilestone2026年7月23日Training Compute-Optimal Large Language Models (Chinchilla)llmscaling-lawmilestone2026年7月23日Character Region Awareness for Text Detection (CRAFT)ocrtext-detectionmilestone2026年7月23日An End-to-End Trainable Neural Network for Image-based Sequence Recognition (CRNN)ocrtext-recognitionmilestone2026年7月23日Denoising Diffusion Probabilistic Models (DDPM)generativediffusioncvmilestone2026年7月23日DiffBrush: Diffusion Brush for Handwritten Text-Line Generationhtghandwritingdiffusiontext-linestyle-disentanglesynthetic-dataocr2026年7月23日DiffusionPen: Towards Controlling the Style of Handwritten Text Generationhtghandwritinglatent-diffusionfew-shotstylehtrsynthetic-data2026年7月23日DiT: Self-supervised Pre-training for Document Image Transformerocrdocument-aiself-supervisedpretrainingmilestone2026年7月23日Dolphin-v2: Scalable Anchor Prompting for Document Parsingocrdocument-parsinganchor-promptingphotographed-docsfine-grained-detectionhybrid-parsingfrontend2026年7月23日Dolphin: Document Image Parsing via Heterogeneous Anchor Promptingocrdocument-parsinganchor-promptingparallel-decodingtwo-stageanalyze-then-parsefrontendbackend2026年7月23日OCR-free Document Understanding Transformer (Donut)ocrdocument-understandingocr-freemilestone2026年7月23日Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networkscvdetectionrpnanchormilestone2026年7月23日GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documentsocrbenchmarkdocument-reasoningmultimodalmilestone2026年7月23日Going Deeper with Convolutions (GoogLeNet / Inception v1)cnncvimagenetinceptionmilestone2026年7月23日General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model (GOT)ocrocr-20unifiedend-to-endmilestone2026年7月23日Language Models are Few-Shot Learners (GPT-3)llmnlpscalingmilestone2026年7月23日HunyuanOCR: Commercial-grade Lightweight 1B OCR VLMocrvlmsmall-modelunifiedspottingtranslationend-to-end2026年7月23日Training LMs to Follow Instructions with Human Feedback (InstructGPT/RLHF)llmrlhfalignmentmilestone2026年7月23日Agentic Document Extraction Gen2 (DPT-3 family)document-aicommerciallandingaidpt3ocr2026年7月23日Detecting Signatures, Stamps and Seals (ADE Attestation Detection)document-ailandingaiattestationocr2026年7月23日LayoutLMv3: Pre-training for Document AI with Unified Text and Image Maskingocrdocument-understandinglayoutpretrainingmilestone2026年7月23日LLaMA: Open and Efficient Foundation Language Modelsllmopen-sourcemilestone2026年7月23日Logics-Parsing-Omni: Unified Omni Parsing (documents/images/AV)ocrdocument-parsingomnievidence-anchoringprogressive-parsingreading-order2026年7月23日Logics-Parsing: End-to-End LVLM with RL for Layout & Reading Orderocrdocument-parsingreinforcement-learningreading-orderlayouthtml-outputend-to-end2026年7月23日Semi-Supervised Adaptation of Diffusion Models for HTG (MAE style condition)htghandwritingdiffusionmaesemi-supervisedunseen-stylesynthetic-data2026年7月23日Mask R-CNNcvinstance-segmentationdetectionroialignmilestone2026年7月23日MinerU2.5-Pro: Data-Driven Document Parsing (same 1.2B architecture)ocrdocument-parsingdata-enginetraining-strategygrpoomnidocbenchmilestone2026年7月23日MinerU2.5: A Decoupled VLM for Efficient High-Resolution Document Parsingocrdocument-parsingvlmdecoupledhigh-resolutioncoarse-to-finefrontendbackend2026年7月23日OCRVerse: Towards Holistic OCR in End-to-Endocrholistic-ocrend-to-enddata-engineeringsft-rltext-centricvision-centric2026年7月23日One-DM: One-Shot Diffusion Mimicker for Handwritten Text Generationhtghandwritingdiffusionone-shotstylesynthetic-data2026年7月23日PaddleOCR-VL: 0.9B Ultra-Compact VLM for Document Parsingocrdocument-parsingvlmsmall-modellayoutomnidocbenchfrontendbackend2026年7月23日Phi-4-Mini / Phi-4-Multimodal Technical Report (Mixture-of-LoRAs)llmmultimodalsmall-modelloraphispeechmilestone2026年7月23日Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understandingocrvisual-languagescreenshotpretrainingmilestone2026年7月23日POINTS-Reader: Distillation-Free Adaptation of VLMs for Document Conversionocrdocument-parsingdistillation-freesynthetic-dataself-improvementend-to-end2026年7月23日Qianfan-OCR: Unified End-to-End Document Intelligence (Layout-as-Thought)ocrdocument-intelligenceend-to-endlayout-as-thoughtthinkingreading-orderomnidocbench2026年7月23日Qwen3.5-Omni Technical Reportllmomnimodalmoespeechagentqwenmilestone2026年7月23日Qwen3 Technical Reportllmmoereasoningmultilingualqwenmilestone2026年7月23日Deep Residual Learning for Image Recognition (ResNet)resnetcnncvfoundationmilestone2026年7月23日Sequence to Sequence Learning with Neural Networks (seq2seq)nlpseq2seqmilestone2026年7月23日StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generationhtghandwritingdiffusioncross-lingualimage-to-imagesynthetic-data2026年7月23日Exploring the Limits of Transfer Learning (T5)llmnlptransfermilestone2026年7月23日Textbooks Are All You Need (phi-1)llmdata-qualityphismall-modelmilestone2026年7月23日TrOCR: Transformer-based Optical Character Recognition with Pre-trained Modelsocrtext-recognitiontransformermilestone2026年7月23日U-Net: Convolutional Networks for Biomedical Image Segmentationcvsegmentationbiomedicalencoder-decodermilestone2026年7月23日UniRec-0.1B: Unified Text and Formula Recognition (0.1B)ocrrecognitiontiny-modelhierarchical-supervisionsemantic-decoupled-tokenizerformulatablebackend2026年7月23日VATr: Handwritten Text Generation from Visual Archetypeshtghandwritingtransformerfew-shotvisual-archetyperare-charactersmilestone2026年7月23日Very Deep Convolutional Networks for Large-Scale Image Recognition (VGG)cnncvimagenetdepthmilestone2026年7月23日Efficient Estimation of Word Representations (Word2Vec)nlpembeddingmilestone2026年7月23日You Only Look Once: Unified, Real-Time Object Detection (YOLO)cvdetectionreal-timeone-stagemilestone2026年7月23日Youtu-Parsing: High-Parallelism Decoding for Document Parsingocrdocument-parsingparallel-decodingtoken-parallelismquery-parallelismregion-promptbackendfrontend