BMD-45 is a new large-scale CCTV vehicle detection dataset from developing cities that reveals a 2.5x performance gap for models adapted from prior benchmarks.
PP-OCRv3: More attempts for the improvement of ultra lightweight OCR system
12 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
method 1polarities
use method 1representative citing papers
MNAFT identifies language-agnostic and language-specific neurons via activation analysis and selectively fine-tunes only relevant ones in MLLMs to close the modality gap and outperform full fine-tuning and other methods on image translation benchmarks.
CAPED reduces incidental visual privacy leakage in mobile GUI agents from 0.766 to 0.268 on seeded AndroidWorld tasks by selectively exposing only task-relevant screen content.
StyleTextGen proposes a dual-branch style encoder, text style consistency loss, and mask-guided inference to achieve superior style consistency and cross-lingual performance in multilingual scene text generation on a new bilingual benchmark.
TextDS uses a data-efficient dual-encoder with SWLoRA and CSF to achieve competitive scene text detection robustness under distribution shifts and adverse conditions using 4.9M trainable parameters.
VaaWIT proposes DSAM and VAA modules to adapt LLMs for multilingual web image translation, claiming outperformance over open-source baselines on benchmarks.
CogVLM2 family achieves state-of-the-art results on image and video understanding benchmarks through improved visual expert architecture, higher resolution inputs, and automated temporal grounding for videos.
ASASR recasts generative SR flow into Sobolev Riemannian geometry via colored noise kernels and a Riesz-based parametric adversary to optimize along plausible structural failure tangents, claiming better spectral consistency than baselines.
A proactive EMR assistant using streaming ASR and belief stabilization reaches 0.84 state-event F1, 0.87 retrieval Recall@5, and 83.3% coverage in a controlled pilot of ten doctor-patient dialogues.
PaddleOCR 3.0 releases compact open-source models for OCR, document structure parsing, and information extraction that rival billion-parameter VLMs.
InternVL 1.5 narrows the performance gap to proprietary multimodal models via a stronger transferable vision encoder, dynamic high-resolution tiling, and curated English-Chinese training data.
PP-OCRv6 introduces three tiers of lightweight OCR models (1.5M–34.5M parameters) built on unified MetaFormer blocks with reparameterization that claim superior accuracy to PP-OCRv5 and billion-scale VLMs on in-house benchmarks.
citing papers explorer
-
BMD-45: A Large-Scale CCTV Vehicle Detection Dataset for Urban Traffic in Developing Cities
BMD-45 is a new large-scale CCTV vehicle detection dataset from developing cities that reveals a 2.5x performance gap for models adapted from prior benchmarks.
-
MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation
MNAFT identifies language-agnostic and language-specific neurons via activation analysis and selectively fine-tunes only relevant ones in MLLMs to close the modality gap and outperform full fine-tuning and other methods on image translation benchmarks.
-
CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents
CAPED reduces incidental visual privacy leakage in mobile GUI agents from 0.766 to 0.268 on seeded AndroidWorld tasks by selectively exposing only task-relevant screen content.
-
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
StyleTextGen proposes a dual-branch style encoder, text style consistency loss, and mask-guided inference to achieve superior style consistency and cross-lingual performance in multilingual scene text generation on a new bilingual benchmark.
-
TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts
TextDS uses a data-efficient dual-encoder with SWLoRA and CSF to achieve competitive scene text detection robustness under distribution shifts and adverse conditions using 4.9M trainable parameters.
-
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
VaaWIT proposes DSAM and VAA modules to adapt LLMs for multilingual web image translation, claiming outperformance over open-source baselines on benchmarks.
-
CogVLM2: Visual Language Models for Image and Video Understanding
CogVLM2 family achieves state-of-the-art results on image and video understanding benchmarks through improved visual expert architecture, higher resolution inputs, and automated temporal grounding for videos.
-
Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution
ASASR recasts generative SR flow into Sobolev Riemannian geometry via colored noise kernels and a Riesz-based parametric adversary to optimize along plausible structural failure tangents, claiming better spectral consistency than baselines.
-
A Proactive EMR Assistant for Doctor-Patient Dialogue: Streaming ASR, Belief Stabilization, and Preliminary Controlled Evaluation
A proactive EMR assistant using streaming ASR and belief stabilization reaches 0.84 state-event F1, 0.87 retrieval Recall@5, and 83.3% coverage in a controlled pilot of ten doctor-patient dialogues.
-
PaddleOCR 3.0 Technical Report
PaddleOCR 3.0 releases compact open-source models for OCR, document structure parsing, and information extraction that rival billion-parameter VLMs.
-
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
InternVL 1.5 narrows the performance gap to proprietary multimodal models via a stronger transferable vision encoder, dynamic high-resolution tiling, and curated English-Chinese training data.
-
PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks
PP-OCRv6 introduces three tiers of lightweight OCR models (1.5M–34.5M parameters) built on unified MetaFormer blocks with reparameterization that claim superior accuracy to PP-OCRv5 and billion-scale VLMs on in-house benchmarks.