46 items with this tag.2026年8月13日Muon: MomentUm Orthogonalized by Newton-Schulzoptimizertrainingorthogonalizationnewton-schulzspectralmilestoneinfra2026年8月06日Learning Transferable Visual Models From NL Supervision (CLIP)multimodalcvnlpmilestone2026年8月06日Deformable DETR: Deformable Transformers for End-to-End Object Detectiondetectiondetrdeformable-attentionmulti-scalefrontendmilestone2026年8月06日End-to-End Object Detection with Transformers (DETR)detectiondetrtransformerset-predictionmilestonefrontend2026年8月06日DINO: DETR with Improved DeNoising Anchor Boxesdetectiondetrdenoisingquery-selectionsotafrontendmilestone2026年8月06日Masked Autoencoders Are Scalable Vision Learners (MAE)sslcvtransformermilestone2026年8月06日An Image is Worth 16x16 Words (Vision Transformer, ViT)transformercvmilestone2026年7月29日DiG: Reading and Writing — Discriminative and Generative Modeling for Self-Supervised Text Recognitionocrtext-recognitionself-supervisedmimcontrastive-learningmilestone2026年7月25日ImageNet Classification with Deep CNNs (AlexNet)cnncvimagenetmilestone2026年7月25日DeepSeek-OCR: Contexts Optical Compressionocroptical-compressionvision-tokenmoemilestone2026年7月25日GLM-OCR Technical Reportocrdocument-parsingcompactmulti-token-predictionmilestone2026年7月25日MonkeyOCRv2: A Visual-Text Foundation Model for Document AIocrdocument-aivisual-encoderfoundation-modelmilestone2026年7月25日Nougat: Neural Optical Understanding for Academic Documentsdocument-aiocrformulamilestone2026年7月25日Ovis-OCR2: 0.8B End-to-End Document Parsing (OmniDocBench SOTA)ocrdocument-parsingend-to-endsmall-modeldata-enginesynthetic-datadistillationmilestone2026年7月23日Attention Is All You Need (Transformer)transformerattentionnlpfoundationmilestone2026年7月23日BERT: Pre-training of Deep Bidirectional Transformersnlptransformerpretrainingmilestone2026年7月23日Training Compute-Optimal Large Language Models (Chinchilla)llmscaling-lawmilestone2026年7月23日Character Region Awareness for Text Detection (CRAFT)ocrtext-detectionmilestone2026年7月23日An End-to-End Trainable Neural Network for Image-based Sequence Recognition (CRNN)ocrtext-recognitionmilestone2026年7月23日Denoising Diffusion Probabilistic Models (DDPM)generativediffusioncvmilestone2026年7月23日DiT: Self-supervised Pre-training for Document Image Transformerocrdocument-aiself-supervisedpretrainingmilestone2026年7月23日OCR-free Document Understanding Transformer (Donut)ocrdocument-understandingocr-freemilestone2026年7月23日Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networkscvdetectionrpnanchormilestone2026年7月23日GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documentsocrbenchmarkdocument-reasoningmultimodalmilestone2026年7月23日Going Deeper with Convolutions (GoogLeNet / Inception v1)cnncvimagenetinceptionmilestone2026年7月23日General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model (GOT)ocrocr-20unifiedend-to-endmilestone2026年7月23日Language Models are Few-Shot Learners (GPT-3)llmnlpscalingmilestone2026年7月23日Training LMs to Follow Instructions with Human Feedback (InstructGPT/RLHF)llmrlhfalignmentmilestone2026年7月23日LayoutLMv3: Pre-training for Document AI with Unified Text and Image Maskingocrdocument-understandinglayoutpretrainingmilestone2026年7月23日LLaMA: Open and Efficient Foundation Language Modelsllmopen-sourcemilestone2026年7月23日Mask R-CNNcvinstance-segmentationdetectionroialignmilestone2026年7月23日MinerU2.5-Pro: Data-Driven Document Parsing (same 1.2B architecture)ocrdocument-parsingdata-enginetraining-strategygrpoomnidocbenchmilestone2026年7月23日Phi-4-Mini / Phi-4-Multimodal Technical Report (Mixture-of-LoRAs)llmmultimodalsmall-modelloraphispeechmilestone2026年7月23日Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understandingocrvisual-languagescreenshotpretrainingmilestone2026年7月23日Qwen3.5-Omni Technical Reportllmomnimodalmoespeechagentqwenmilestone2026年7月23日Qwen3 Technical Reportllmmoereasoningmultilingualqwenmilestone2026年7月23日Deep Residual Learning for Image Recognition (ResNet)resnetcnncvfoundationmilestone2026年7月23日Sequence to Sequence Learning with Neural Networks (seq2seq)nlpseq2seqmilestone2026年7月23日Exploring the Limits of Transfer Learning (T5)llmnlptransfermilestone2026年7月23日Textbooks Are All You Need (phi-1)llmdata-qualityphismall-modelmilestone2026年7月23日TrOCR: Transformer-based Optical Character Recognition with Pre-trained Modelsocrtext-recognitiontransformermilestone2026年7月23日U-Net: Convolutional Networks for Biomedical Image Segmentationcvsegmentationbiomedicalencoder-decodermilestone2026年7月23日VATr: Handwritten Text Generation from Visual Archetypeshtghandwritingtransformerfew-shotvisual-archetyperare-charactersmilestone2026年7月23日Very Deep Convolutional Networks for Large-Scale Image Recognition (VGG)cnncvimagenetdepthmilestone2026年7月23日Efficient Estimation of Word Representations (Word2Vec)nlpembeddingmilestone2026年7月23日You Only Look Once: Unified, Real-Time Object Detection (YOLO)cvdetectionreal-timeone-stagemilestone