AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,965Total entries
1Added by human
86,964Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

Can LLMs Be Constrained to the Past? Improving Knowledge Cutoff through Recall-Based Prompting

DGX agent

arXiv:2606.05804v1 Announce Type: new Abstract: Prompted knowledge cutoff instructs a large language model (LLM) to act as if information beyond a specified cutoff date were unavailable. However, prio

model-releasesarxiv-cs-cl
5 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation

DGX agent

arXiv:2606.05988v1 Announce Type: cross Abstract: Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs. We study post-hoc compression

model-releasesarxiv-cs-cl
5 Jun 2026
Research

From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment

DGX agent

arXiv:2606.05180v1 Announce Type: new Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts,

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

DGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

DGX agent

arXiv:2606.06197v1 Announce Type: new Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face challenges in a

model-releasesarxiv-cs-cl
5 Jun 2026
Research

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

DGX agent

arXiv:2606.06286v1 Announce Type: new Abstract: Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather th

researcharxiv-cs-cl
5 Jun 2026
Model Releases

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

DGX agent

arXiv:2606.05677v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon ta

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

DGX agent

arXiv:2606.05176v1 Announce Type: new Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-s

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces

DGX agent

arXiv:2606.05402v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluat

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework

DGX agent

arXiv:2606.05570v1 Announce Type: new Abstract: Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often invol

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

DGX agent

arXiv:2606.06476v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning

DGX agent

arXiv:2606.05576v1 Announce Type: new Abstract: Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images -

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

DGX agent

arXiv:2606.04291v1 Announce Type: new Abstract: 3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains f

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

Adaptive Minds: Empowering Agents with LoRA-as-Tools

DGX agent

arXiv:2510.15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke. We hyp

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers

DGX agent

arXiv:2606.04752v1 Announce Type: cross Abstract: Transformers consuming multi-channel scalar signals must embed C simultaneous values into one d_{ext{model}}-dimensional vector per time step. We empi

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Bypassing Prompt Guards in Production with Controlled-Release Prompting

DGX agent

arXiv:2510.01529v3 Announce Type: replace Abstract: Ball et al. recently established that prompt filtering for AI alignment faces a fundamental barrier: under standard cryptographic assumptions, no fi

model-releasesarxiv-cs-lg
4 Jun 2026
Model Releases

Can I Take Another Dose? Evaluating LLM Decision-Making Under Temporal Uncertainty in OTC Dosing QA

DGX agent

arXiv:2606.04262v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for everyday health questions, including whether a user can safely take another dose of an over-the

model-releasesarxiv-cs-ai
4 Jun 2026
Research

Enhancing Hallucination Detection through Noise Injection

DGX agent

arXiv:2502.03799v4 Announce Type: replace Abstract: Large Language Models (LLMs) are prone to generating plausible yet incorrect responses, known as hallucinations. Effectively detecting hallucination

researcharxiv-cs-cl
4 Jun 2026
Agents

FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games

DGX agent

arXiv:2606.04751v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks. Yet whether these systems can effectively engage in for

agentsarxiv-cs-ai
4 Jun 2026
Safety

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

DGX agent

arXiv:2603.07445v2 Announce Type: replace Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when

safetyarxiv-cs-cl
4 Jun 2026
Model Releases

Knowledge Index of Noah's Ark

DGX agent

arXiv:2606.05104v1 Announce Type: new Abstract: Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotat

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

LLMs + Persona-Plug = Personalized LLMs

DGX agent

arXiv:2409.11901v2 Announce Type: replace Abstract: Personalization plays a critical role in numerous language tasks and applications, since users with the same requirements may prefer diverse outputs

model-releasesarxiv-cs-cl
4 Jun 2026
Research

Making Expert Reasoning Learnable with Self-Distillation

DGX agent

arXiv:2602.02405v2 Announce Type: replace-cross Abstract: Improving the reasoning capabilities of large language models (LLMs) typically relies either on the model's ability to sample a correct soluti

researcharxiv-cs-ai
4 Jun 2026
Model Releases

Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge

DGX agent

arXiv:2606.04581v1 Announce Type: cross Abstract: Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propos

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

SAM 3D: 3Dfy Anything in Images

DGX agent

arXiv:2511.16624v2 Announce Type: replace-cross Abstract: We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single i

model-releasesarxiv-cs-ai
4 Jun 2026
Research

Spectral Scaling Laws of Muon

DGX agent

arXiv:2606.04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-th

researcharxiv-cs-ai
4 Jun 2026
Model Releases

VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

DGX agent

arXiv:2606.04244v1 Announce Type: new Abstract: Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a proble

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

DGX agent

arXiv:2606.04588v1 Announce Type: new Abstract: Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide lim

model-releasesarxiv-cs-cl
4 Jun 2026
Research

AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making

DGX agent

arXiv:2606.03198v1 Announce Type: cross Abstract: Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluatio

researcharxiv-cs-ai
3 Jun 2026
Model Releases

An Asymptotic Theory of Chain-of-Thought in In-Context Learning

DGX agent

arXiv:2606.03217v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermed

model-releasesarxiv-cs-lg
3 Jun 2026
Research

Anomalies in Multivariate Time Series Benchmarks Are Mostly Univariate

DGX agent

arXiv:2606.02670v1 Announce Type: cross Abstract: Many recent multivariate time series anomaly detection (MT-SAD) models incorporate cross-channel modeling, under the implicit assumption that the stru

researcharxiv-cs-ai
3 Jun 2026
Model Releases

ATLAS: A Large-Scale Evaluation Benchmark for Adversarial LiDAR Perception

DGX agent

arXiv:2606.02924v1 Announce Type: new Abstract: Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and pot

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification

DGX agent

arXiv:2606.03031v1 Announce Type: new Abstract: Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs

DGX agent

arXiv:2606.03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Causal Neural Probabilistic Circuits

DGX agent

arXiv:2603.01372v2 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) enhance the interpretability of end-to-end neural networks by introducing a layer of concepts and predicting

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

DECA: Decentralizing Block-Wise Adam for Efficient LLM Full-Parameter Fine-Tuning on Non-IID Data

DGX agent

arXiv:2606.03209v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) in privacy-sensitive and resource-constrained environments remains challenging. Since training data are often d

model-releasesarxiv-cs-lg
3 Jun 2026
Research

DiffUNet^2: Bidirectional Prediction, Probabilistic Generation and Collaborative Visual Discovery for Scientific Data

DGX agent

arXiv:2606.03926v1 Announce Type: cross Abstract: Modeling temporal evolution is important to analyzing and reasoning about scientific phenomena, yet most machine learning methods provide deterministi

researcharxiv-cs-lg
3 Jun 2026
Safety

Effect of Demographic Bias on Skin Lesion Classification

DGX agent

arXiv:2606.03214v1 Announce Type: new Abstract: In this study, we evaluate the performance of skin lesion classification using ResNet-based convolutional models, focusing on the impact of demographic

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

From Script to Semantics: Prompting Strategies for African NLI

DGX agent

arXiv:2606.03304v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

DGX agent

arXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

DGX agent

arXiv:2602.06960v3 Announce Type: replace-cross Abstract: Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, c

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Let the Dynamics Flow: Stable Flow Matching Dynamical Systems

DGX agent

arXiv:2606.03834v1 Announce Type: new Abstract: Flow matching has recently emerged as a powerful approach for imitation learning, enabling scalable, expressive, and multimodal motion policies. However

model-releasesarxiv-cs-ro
3 Jun 2026
Model Releases

Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility

DGX agent

arXiv:2606.03291v1 Announce Type: new Abstract: Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. Howev

model-releasesarxiv-cs-cl
3 Jun 2026
Agents

MUSE: A Unified Agentic Harness for MLLMs

DGX agent

arXiv:2606.03005v1 Announce Type: cross Abstract: Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze fr

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

SCOPE: Real-Time Natural Language Camera Agent at the Edge

DGX agent

arXiv:2606.02951v1 Announce Type: cross Abstract: Deploying language-driven agents in robotics requires evaluations that reflect real-world task demands: natural-language instructions with reproducibl

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

DGX agent

arXiv:2606.03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for de

model-releasesarxiv-cs-ai
3 Jun 2026
Research

A Direct Approach for Handling Contextual Bandits with Latent State Dynamics

DGX agent

arXiv:2604.08149v2 Announce Type: replace Abstract: We consider a linear contextual bandit model where contexts and rewards are governed by a finite hidden Markov chain. We first revisit the simplifie

researcharxiv-cs-lg
2 Jun 2026
Model Releases

A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL

DGX agent

arXiv:2606.02398v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation,

model-releasesarxiv-cs-cl
2 Jun 2026
← Previous
1…316317318319320…1065
Next →