AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs

DGX agent

arXiv:2606.06203v1 Announce Type: new Abstract: Input length and the position of relevant information are widely cited as the primary causes of degraded LLM long-context performance. Here, we study le

model-releasesarxiv-cs-cl
5 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments

DGX agent

arXiv:2606.06217v1 Announce Type: new Abstract: When a disaster unfolds, responders must answer not only what is happening, but also why it is happening, what will happen next, and what to do now, oft

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding

DGX agent

arXiv:2505.05026v5 Announce Type: replace Abstract: User interface (UI) design goes beyond visuals to shape user experience (UX), underscoring the shift toward UI/UX as a unified concept. While recent

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections

DGX agent

arXiv:2508.15851v2 Announce Type: replace Abstract: Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs

DGX agent

arXiv:2606.05569v1 Announce Type: new Abstract: Mispronunciation Detection and Diagnosis (MDD) has gained increasing importance in computer-assisted language learning and speech technology in recent y

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

DGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving

DGX agent

arXiv:2601.21288v2 Announce Type: replace-cross Abstract: Autonomous driving is an important and safety-critical task, and recent advances in LLMs/VLMs have opened new possibilities for reasoning and

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Efficient Punctuation Restoration via Weighted Lookahead Scoring Method for Streaming ASR Systems

DGX agent

arXiv:2606.05179v1 Announce Type: new Abstract: Punctuation restoration improves ASR (Automatic Speech Recognition) readability. However streaming ASR requires online decisions with limited future con

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge

DGX agent

arXiv:2605.24500v2 Announce Type: replace Abstract: This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC V

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

English-to-Prakrit Machine Translation via Multilingual Transfer Learning

DGX agent

arXiv:2606.06038v1 Announce Type: new Abstract: We study English-to-Prakrit machine translation in a low-resource setting where the target language is unsupported by IndicTrans2. We adapt the multilin

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics

DGX agent

arXiv:2606.05168v1 Announce Type: new Abstract: Training on synthetic data causes model collapse, but existing analyses treat this as single-chain degradation. In reality, the AI ecosystem involves cr

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models

DGX agent

arXiv:2606.05874v1 Announce Type: new Abstract: Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-n

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Facial-R1: Aligning Reasoning and Recognition for Facial Emotion Analysis

DGX agent

arXiv:2511.10254v2 Announce Type: replace Abstract: Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrat

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models

DGX agent

arXiv:2606.05949v1 Announce Type: new Abstract: Scientific illustrations are essential tools for communicating research findings, especially in natural science, where they visualize complex concepts a

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

FATE: Focal-modulated Attention Encoder for Multivariate Time-series Forecasting

DGX agent

arXiv:2408.11336v3 Announce Type: replace-cross Abstract: Climate change stands as one of the most pressing global challenges of the twenty-first century, with far-reaching consequences such as rising

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

DGX agent

arXiv:2606.06211v1 Announce Type: new Abstract: Automatic speech recognition (ASR) has advanced remarkably for standard speech; however, pathological speech from neurological conditions remains a sign

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

From Self to Other: Evaluating Demographic Perspective-Taking in LLM Hate Speech Annotation

DGX agent

arXiv:2606.06266v1 Announce Type: new Abstract: Hate speech detection is inherently subjective: people from different demographic groups perceive the same content very differently. Collecting enough a

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery

DGX agent

arXiv:2602.19190v4 Announce Type: replace Abstract: Research on the intelligent interpretation of all-weather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applicati

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Generic Triple-Latent Compression with Gated Associative Retrieval

DGX agent

arXiv:2606.05175v1 Announce Type: new Abstract: We study generic triple-latent sequence models that maintain a running token state and compressed pair-memory pathway to capture higher-order token inte

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework

DGX agent

arXiv:2603.08491v2 Announce Type: replace Abstract: Cross-modal Geo-localization (CMGL) matches ground-level text descriptions with geo-tagged aerial imagery, which is crucial for pedestrian navigatio

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

DGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Harnessing Structural Context for Entity Alignment Foundation Models

DGX agent

arXiv:2606.06109v1 Announce Type: new Abstract: Entity alignment (EA) aims to identify equivalent entities across heterogeneous knowledge graphs (KGs) and is a key component of knowledge fusion and cr

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

DGX agent

arXiv:2601.02730v3 Announce Type: replace Abstract: Visual localization on standard-definition (SD) maps has emerged as a promising low-cost and scalable solution for autonomous driving. However, exis

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

DGX agent

arXiv:2606.06388v1 Announce Type: cross Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly pos

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

IA-RAG: Interval-Algebra-Driven Temporal Reasoning for Dynamic Knowledge Retrieval

DGX agent

arXiv:2606.06044v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has shown strong effectiveness in grounding Large Language Models (LLMs) with external knowledge. However, existing

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

DGX agent

arXiv:2606.06197v1 Announce Type: new Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face challenges in a

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

DGX agent

arXiv:2606.05172v1 Announce Type: cross Abstract: Diffusion-based image editing has achieved strong visual fidelity under natural language instructions, yet most existing systems still operate at the

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

KV-Control: Parameter-Efficient K/V Injection for Trajectory-Controlled Text-to-Motion

DGX agent

arXiv:2606.05624v1 Announce Type: new Abstract: Text-conditioned 3D human motion models now synthesize plausible motions from prompts, but practical animation and embodied-agent workflows rarely stop

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

DGX agent

arXiv:2606.06087v1 Announce Type: new Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substa

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Less is MoE: Trimming Experts in Domain-Specialist Language Models

DGX agent

arXiv:2606.05538v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models achieve strong performance through conditional computation, but their large parameter footprint poses deployment chall

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

LiAuto-GeoX: Efficient Grounded Driving Transformer

DGX agent

arXiv:2606.05774v1 Announce Type: new Abstract: Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for auton

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

LightVesselNet: An Ultra-Lightweight Sub-100K Parameter Network for Retinal Blood Vessel Segmentation

DGX agent

arXiv:2606.05354v1 Announce Type: new Abstract: Retinal blood vessel segmentation plays a vital role in the early detection of diabetic retinopathy and glaucoma. While recent deep learning models have

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

LLM-Guided ANN Index Optimization for Human-Object Interaction Retrieval

DGX agent

arXiv:2606.05489v1 Announce Type: new Abstract: Retrieval systems underpin modern AI applications -- spanning visual search, recommendation engines, and multi-modal question answering. Modern multi-st

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

DGX agent

arXiv:2606.05486v1 Announce Type: new Abstract: Prompt ambiguity is a common source of failure in large language models, but is difficult to localize because it is a latent property of the prompt, whi

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

DGX agent

arXiv:2606.05677v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon ta

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

DGX agent

arXiv:2606.06042v1 Announce Type: new Abstract: Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier fie

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

LoRi: Low-Rank Distillation for Implicit Reasoning

DGX agent

arXiv:2606.05315v1 Announce Type: new Abstract: Implicit chain-of-thought (iCoT) methods aim to internalize reasoning in large language models, but often underperform explicit CoT prompting. We empiri

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

MAviS: A Multimodal Conversational Assistant For Avian Species

DGX agent

arXiv:2603.07294v2 Announce Type: replace Abstract: Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monit

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

DGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following

DGX agent

arXiv:2606.06058v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards is ideal for multi-constraint instruction following, yet standard group-relative policy optimization (G

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Noise-Aware Visual Representation Learning for Medical Visual Question Answering

DGX agent

arXiv:2606.05535v1 Announce Type: new Abstract: Medical visual question answering (Med-VQA) has strong potential for clinical decision support by enabling AI models to interpret medical images and ans

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Oklch+: A Three-Parameter Extension of Oklab for Improved Color Difference Prediction

DGX agent

arXiv:2606.05255v1 Announce Type: cross Abstract: Oklab and its cylindrical representation Oklch are widely adopted in interpolation and design workflows as perceptually motivated color spaces, but th

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

OLIVE: Online Low-Rank Incremental Learning for Efficient Adaptive Exoskeletons

DGX agent

arXiv:2606.05234v1 Announce Type: new Abstract: Wearable exoskeleton systems hold promise for restoring mobility in individuals with physical impairments, yet most existing controllers rely on static

model-releasesarxiv-cs-ro
5 Jun 2026
Model Releases

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

DGX agent

arXiv:2606.06481v1 Announce Type: new Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-writt

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

DGX agent

arXiv:2606.05176v1 Announce Type: new Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-s

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Physics-Guided Deep Unfolding for Blind Cross-Sensor Spectral Super-Resolution via Learning the Spectral Transformation Function

DGX agent

arXiv:2606.05759v1 Announce Type: new Abstract: Hyperspectral imaging provides rich spectral information for quantitative remote sensing, yet hyperspectral sensors remain costly and thus unavailable i

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

DGX agent

arXiv:2606.05744v1 Announce Type: new Abstract: Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Predict and Reconstruct: Joint Objectives for Self-Supervised Language Representation Learning

DGX agent

arXiv:2606.05173v1 Announce Type: new Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are st

model-releasesarxiv-cs-cl
5 Jun 2026
← Previous
1…152153154155156…361
Next →