AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
50,764 results
Model Releases

PreAct-Bench: Benchmarking Predictive Monitoring in LLMs

DGX agent

arXiv:2606.09890v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objecti

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Quo Vadis, Visual In-Context Learning? A Unified Benchmark Across Domains and Tasks

DGX agent

arXiv:2606.10967v1 Announce Type: new Abstract: Visual in-context learning has been proposed as a pathway towards dynamic models that can generate predictions based on a provided context and thereby c

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning

DGX agent

arXiv:2510.14828v3 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

SAFE: An LLM-as-Verifier Framework for Evidence-Grounded Multi-Hop Reasoning

DGX agent

arXiv:2604.01993v2 Announce Type: replace-cross Abstract: Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, where models reach correct answers through invalid

model-releasesarxiv-cs-ai
10 Jun 2026
Applications

Segmentation-Driven Monocular Shape from Polarization based on Physical Model

DGX agent

arXiv:2601.04776v2 Announce Type: replace Abstract: Monocular shape-from-polarization (SfP) leverages the intrinsic relationship between light polarization properties and surface geometry to recover s

applicationsarxiv-cs-cv
10 Jun 2026
Research

Superficial Beliefs in LLM Decision-Making

DGX agent

arXiv:2606.11016v1 Announce Type: new Abstract: We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic u

researcharxiv-cs-ai
10 Jun 2026
Model Releases

U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training

DGX agent

arXiv:2606.11032v1 Announce Type: new Abstract: Existing deep learning models for Positron Emission Tomography (PET) image denoising often suffer from severe performance degradation under distribution

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

Validation-Stage Combinatorial Fusion Analysis for Imbalanced Credit-Card Fraud Detection

DGX agent

arXiv:2606.10393v1 Announce Type: new Abstract: Credit-card fraud detection is difficult because fraudulent transactions are rare, costly, and unevenly distributed. Strong gradient-boosted tree models

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

WebChallenger: A Reliable and Efficient Generalist Web Agent

DGX agent

arXiv:2606.10423v1 Announce Type: new Abstract: Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents

DGX agent

arXiv:2606.11045v1 Announce Type: new Abstract: Reusing a held-out benchmark adaptively should, in principle, invite overfitting. Yet benchmark-driven machine learning (ML) has produced surprisingly l

model-releasesarxiv-cs-ai
10 Jun 2026
Research

A Machine Learning-Enhanced Hopf-Cole Formulation for Nonlinear Gas Flow in Porous Media

DGX agent

arXiv:2603.11250v2 Announce Type: replace-cross Abstract: Accurate modeling of gas flow through porous media is critical for many technological applications, including reservoir performance prediction

researcharxiv-cs-lg
9 Jun 2026
Safety

An Agency-Transferring Model-Free Policy Enhancement Technique

DGX agent

arXiv:2606.09825v1 Announce Type: cross Abstract: Training reinforcement learning (RL) policies from scratch is costly: it requires careful reward and environment design, extensive tuning, and substan

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

DGX agent

arXiv:2606.07643v1 Announce Type: cross Abstract: Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their a

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs

DGX agent

arXiv:2606.08840v1 Announce Type: new Abstract: Code generation models are typically compared using compact execution benchmarks and aggregate pass rates, but such summaries obscure how performance va

model-releasesarxiv-cs-ai
9 Jun 2026
Research

CRAG: Can 3D Generative Models Help 3D Assembly?

DGX agent

arXiv:2602.22629v2 Announce Type: replace Abstract: Most existing 3D assembly methods treat the problem as pure pose estimation, rearranging observed parts via rigid transformations. In contrast, huma

researcharxiv-cs-cv
9 Jun 2026
Model Releases

CRANE: Knowledge Editing for Reasoning MLLMs

DGX agent

arXiv:2606.09033v1 Announce Type: new Abstract: The emergence of reasoning multimodal large language models (MLLMs), which generate explicit chain-of-thought (CoT) reasoning before producing answers,

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Decomposable Neuro Symbolic Regression

DGX agent

arXiv:2511.04124v3 Announce Type: replace Abstract: Symbolic regression (SR) models complex systems by discovering mathematical expressions that capture underlying relationships in observed data. Howe

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks

DGX agent

arXiv:2606.07970v1 Announce Type: cross Abstract: Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with o

model-releasesarxiv-cs-ai
9 Jun 2026
Agents

EditSSC: Toward Editable Semantic Occupancy Scenes with Unconditional Diffusion Models

DGX agent

arXiv:2606.09273v1 Announce Type: new Abstract: 3D semantic scene generation is crucial for autonomous driving applications, yet most methods rely on complex 3D-specific architectures such as triplane

agentsarxiv-cs-cv
9 Jun 2026
Model Releases

How Much Capacity Does EEG Denoising Need? Ultra-Compact Networks reveal Benchmark Saturation and Metric-Utility Gap

DGX agent

arXiv:2606.08594v1 Announce Type: new Abstract: Deep learning EEG denoising architectures have scaled from tens of thousands to tens of millions of parameters, yet no prior study has isolated model ca

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation

DGX agent

arXiv:2601.04498v2 Announce Type: replace-cross Abstract: Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation

DGX agent

arXiv:2510.20182v2 Announce Type: replace Abstract: Pedestrian simulation traditionally relies on expert-tuned, hand-crafted models that limit scalability and generalization. Meanwhile, large-scale vi

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

Petri Net Modeling and Deadlock-Free Scheduling of Attachable Heterogeneous AGV Systems

DGX agent

arXiv:2508.00724v2 Announce Type: replace-cross Abstract: The increasing demand for flexible automation has accelerated the adoption of heterogeneous automated guided vehicles (AGVs). This work invest

safetyarxiv-cs-ro
9 Jun 2026
Model Releases

PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits

DGX agent

arXiv:2510.17947v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are improving at an exceptional rate. With the advent of agentic workflows, multi-turn dialogue has become the de

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks

DGX agent

arXiv:2606.07968v1 Announce Type: cross Abstract: Reasoning-capable large language models can be induced to spend their generation budget on injected decoy tasks rather than answering the user's quest

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Scaling Laws for Masked-Reconstruction Transformers on Single-Cell Transcriptomics

DGX agent

arXiv:2602.15253v2 Announce Type: replace Abstract: Neural scaling laws -- power-law relationships between loss, model size, and data -- have been extensively documented for language and vision transf

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems

DGX agent

arXiv:2606.08034v1 Announce Type: cross Abstract: Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions. However, existing s

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

DGX agent

arXiv:2606.07608v1 Announce Type: cross Abstract: We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standar

model-releasesarxiv-cs-ai
9 Jun 2026
Research

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

DGX agent

arXiv:2606.09019v1 Announce Type: cross Abstract: Codec-based autoregressive (AR) speech language models have achieved strong text-to-speech (TTS) quality by modeling speech as sequences of discrete a

researcharxiv-cs-ai
9 Jun 2026
Model Releases

VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation

DGX agent

arXiv:2606.07992v1 Announce Type: new Abstract: As the Model Context Protocol (MCP) standardizes tool-calling for autonomous agents, it introduces a critical, unexamined attack surface: the error-hand

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?

DGX agent

arXiv:2606.07872v1 Announce Type: new Abstract: When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual e

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning

DGX agent

arXiv:2602.01357v2 Announce Type: replace Abstract: Self-play post-training methods has emerged as an effective approach for finetuning large language models and turn the weak language model into stro

safetyarxiv-cs-lg
9 Jun 2026
Model Releases

Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline

DGX agent

arXiv:2606.07965v1 Announce Type: new Abstract: Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks. However, the significant differences between industrial and natur

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

ZIPP:Zero-shot Image Personalization from Personas

DGX agent

arXiv:2606.08841v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate a

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

DGX agent

arXiv:2606.06526v1 Announce Type: new Abstract: Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development

DGX agent

arXiv:2606.07207v1 Announce Type: cross Abstract: Confidence-based loss weighting is usually avoided in generative models because it accelerates errors when the model is confidently wrong, but this in

model-releasesarxiv-cs-lg
8 Jun 2026
Research

Explaining Unsupervised Disease Staging in Huntington's Disease: Insights into Model Representations and Clusters

DGX agent

arXiv:2606.07135v1 Announce Type: new Abstract: Huntington's disease (HD) is a progressive neurodegenerative disorder that affects motor, cognitive, and behavioral functions, where accurate characteri

researcharxiv-cs-lg
8 Jun 2026
Model Releases

LiQSS: Post-Transformer Linear Quantum-Inspired State-Space Tensor Networks for Real-Time 6G

DGX agent

arXiv:2601.12375v3 Announce Type: replace-cross Abstract: Proactive and agentic control in Sixth-Generation (6G) Open Radio Access Networks (O-RAN) requires control-grade prediction under stringent Ne

model-releasesarxiv-cs-lg
8 Jun 2026
Tutorials

Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning

DGX agent

arXiv:2602.11201v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) explanations are widely used to interpret how language models solve complex problems, yet it remains unclear whether these st

tutorialsarxiv-cs-cl
8 Jun 2026
Research

Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory

DGX agent

arXiv:2606.06624v1 Announce Type: new Abstract: In the current era of deep learning and especially generative models, there is significant investment in training very large generative models. Thus far

researcharxiv-cs-lg
8 Jun 2026
Model Releases

Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces

DGX agent

arXiv:2606.06941v1 Announce Type: new Abstract: Large language models (LLMs) now solve a wide range of expert-level exams at or above human level, yet remain brittle on specialised, evidence-intensive

model-releasesarxiv-cs-ai
8 Jun 2026
Safety

SafeGene: Reusable Adapters for Transferable Safety Alignment

DGX agent

arXiv:2606.06519v1 Announce Type: new Abstract: Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vul

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors

DGX agent

arXiv:2606.06891v1 Announce Type: new Abstract: Despite advances in 3D scene understanding, existing 3D Large Multimodal Models operate in offline settings, requiring complete scene observations or pr

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

DGX agent

arXiv:2606.07167v1 Announce Type: cross Abstract: Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lack

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

DGX agent

arXiv:2606.06538v1 Announce Type: new Abstract: In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

Can AI Refute Economic Theory? Evidence from Beyond the Knowledge Cutoff

DGX agent

arXiv:2606.05383v1 Announce Type: cross Abstract: Can artificial intelligence (AI) refute economic theory? I document experiments in which I asked several AI models (Gemini, Refine, Claude, and ChatGP

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention

DGX agent

arXiv:2602.13255v2 Announce Type: replace Abstract: We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models. Existing benchmarks measure tas

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions

DGX agent

arXiv:2606.06322v1 Announce Type: new Abstract: GUI agents - vision-based models that control desktops, web browsers, and mobile devices through graphical user interfaces - promise to automate a wide

model-releasesarxiv-cs-ai
6 Jun 2026
← Previous
1…283284285286287…1058
Next →