TechniqueRLHF / Alignment8 recent entries12 Aug 2026MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level AlignmentarXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image repr→12 Aug 2026LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int
TechniqueRAG8 recent entries31 Jul 2026Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal ReasoningarXiv:2508.16129v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although se→31 Jul 2026AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic ReportsarXiv:2601.15297v3 Announce Type: replace Abstract: Reliable question answering over long institutional documents requires more than topical retrieval: a system must localize the exact passage that su→4 Aug 2026Prompt-Driven Simulation with Feature Perturbation for Cross-Domain Few-Shot Object DetectionarXiv:2608.01348v1 Announce Type: new Abstract: Data augmentation, which simulates diverse visual variations to expand the source distribution and induce synthetic domain shifts, is a simple yet effec→10 Aug 2026Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image GenerationarXiv:2608.06751v1 Announce Type: cross Abstract: Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical→11 Aug 2026VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative ScaffoldingarXiv:2608.09698v1 Announce Type: cross Abstract: Great fiction earns its verisimilitude through precise details, from how a longsword is gripped to pierce armor gaps to why a bleeding corpse cannot y→11 Aug 2026RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and EditingarXiv:2608.09186v1 Announce Type: new Abstract: Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geom→12 Aug 2026Situation Graph Prediction for User Perspective ModelingarXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data →12 Aug 2026Order Matters: LVLMs as Judges for Temporal Reasoning in Image SequencesarXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud
TechniqueAgents8 recent entries11 Aug 2026Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMsarXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent →11 Aug 2026ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with CouponsarXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise→12 Aug 2026When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic UnderwritingarXiv:2606.16465v2 Announce Type: replace Abstract: AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or transferred. →12 Aug 2026Situation Graph Prediction for User Perspective ModelingarXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data →12 Aug 2026Order Matters: LVLMs as Judges for Temporal Reasoning in Image SequencesarXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud→12 Aug 2026LLM Agents Factory: Retrieval of Domain-Specific LLM AgentsarXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deploymen→12 Aug 2026Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware ReformulationarXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive fie→12 Aug 2026Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive PerceptionarXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa
TechniqueFine-tuning8 recent entries11 Aug 2026GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory CurriculumarXiv:2603.28533v3 Announce Type: replace Abstract: Agentic knowledge graph question answering (KGQA) requires an agent to iteratively interact with knowledge graphs (KGs), posing challenges in both t→11 Aug 2026EvBS: Event-guided Blur Synthesis for Domain-adaptive Motion DeblurringarXiv:2608.08066v1 Announce Type: new Abstract: Motion deblurring has achieved remarkable progress with deep learning, yet pre-trained deblurring models often suffer from performance degradation in re→11 Aug 2026DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese DialectsarXiv:2608.08067v1 Announce Type: cross Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to→11 Aug 2026AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action ModelsarXiv:2603.05868v2 Announce Type: replace Abstract: Despite remarkable progress in Vision-Language-Action models (VLAs) for robot manipulation, these large pre-trained models require fine-tuning to be→12 Aug 2026LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int→12 Aug 2026LLM Agents Factory: Retrieval of Domain-Specific LLM AgentsarXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deploymen→12 Aug 2026Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware ReformulationarXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive fie→12 Aug 2026CARE: Confidence-Aware Reasoning for Reliable Medical VQAarXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu
TechniqueMultimodal8 recent entries12 Aug 2026Order Matters: LVLMs as Judges for Temporal Reasoning in Image SequencesarXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud→12 Aug 2026MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level AlignmentarXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image repr→12 Aug 2026LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int→12 Aug 2026Longitudinal 3D Foundation Modeling for Neoadjuvant Breast Cancer Response Prediction from Serial DCE-MRIarXiv:2608.09991v1 Announce Type: cross Abstract: Pathologic complete response (pCR) is an important endpoint in neoadjuvant chemotherapy (NAC) for breast cancer, and predicting pCR from imaging durin→12 Aug 2026FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion EditingarXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc→12 Aug 2026Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian SplattingarXiv:2608.10756v1 Announce Type: cross Abstract: Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before ex→12 Aug 2026DreamOmni3: Scribble-based Editing and GenerationarXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text →12 Aug 2026CARE: Confidence-Aware Reasoning for Reliable Medical VQAarXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu
TechniqueSafety8 recent entries12 Aug 2026What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical ResearcharXiv:2608.10431v1 Announce Type: cross Abstract: Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implemen→12 Aug 2026Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video CaptioningarXiv:2608.11013v1 Announce Type: new Abstract: Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, lead→12 Aug 2026The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction DatasetarXiv:2608.10839v1 Announce Type: new Abstract: This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems traine→12 Aug 2026LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int→12 Aug 2026Introspective Attention Modulation for Safe Text-to-Image GenerationarXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr→12 Aug 2026FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion EditingarXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc→12 Aug 2026Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive PerceptionarXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa→12 Aug 2026CARE: Confidence-Aware Reasoning for Reliable Medical VQAarXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu