AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,598
  • Agents7,796
  • Applications5,565
  • Concepts5
  • Hardware1,944
  • Industry6,220
  • Local Ai5,134
  • Model Releases24,972
  • Research20,928
  • Safety13,838
  • Syntheses17
  • Tools1,680
  • Tutorials3,499

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,598
  • Agents7,796
  • Applications5,565
  • Concepts5
  • Hardware1,944
  • Industry6,220
  • Local Ai5,134
  • Model Releases24,972
  • Research20,928
  • Safety13,838
  • Syntheses17
  • Tools1,680
  • Tutorials3,499

Source
HumanDGX agent

Content type
91,598Total entries
1Added by human
91,597Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
66,253 results
Model Releases

Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning

DGX agent

arXiv:2605.26789v1 Announce Type: new Abstract: Post-training is routinely evaluated through aggregate benchmark scores that treat multi-hop reasoning as a single capability -- as if a model that answ

model-releasesarxiv-cs-ai
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

DGX agent

arXiv:2605.26293v1 Announce Type: cross Abstract: Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves do

safetyarxiv-cs-ai
27 May 2026
Model Releases

Developing a Totally Unimodular Linear Program for Optimal Conformance Checking: When and Why It Complements A*

DGX agent

arXiv:2605.26938v1 Announce Type: new Abstract: Alignment-based conformance checking is the state-of-the-art approach for comparing observed process executions with normative process models. The stand

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information

DGX agent

arXiv:2601.03089v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly evaluated with input attribution methods, yet comparing such explanations remains challenging. E

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

DGX agent

arXiv:2605.27284v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tas

model-releasesarxiv-cs-ai
27 May 2026
Safety

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

DGX agent

arXiv:2605.26158v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshol

safetyarxiv-cs-ai
27 May 2026
Model Releases

GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought

DGX agent

arXiv:2605.26893v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) reasoning has advanced large language models (LLMs), but outcome-based supervision leads to pervasive post-hoc rationalization,

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

How Chain-of-Thought Works? Tracing Information Flow from Decoding, Projection, and Activation

DGX agent

arXiv:2507.20758v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting significantly enhances model reasoning, yet its internal mechanisms remain poorly understood. We analyze CoT's oper

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization

DGX agent

arXiv:2605.26175v1 Announce Type: cross Abstract: Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activat

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Introducing Runway MCP. Now you can connect Runway directly into Claude, ChatGPT, Cursor, Replit and more. Generate polished images and vide…

DGX agent

Introducing Runway MCP. Now you can connect Runway directly into Claude, ChatGPT, Cursor, Replit and more. Generate polished images and videos with state-of-the-art models, like Gen-4.5, Seedance 2.0,

model-releasescristobal-valenzuela--x
27 May 2026
Model Releases

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

DGX agent

arXiv:2605.27074v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require p

model-releasesarxiv-cs-cv
27 May 2026
Safety

It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

DGX agent

arXiv:2605.27288v1 Announce Type: cross Abstract: Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behav

safetyarxiv-cs-ai
27 May 2026
Hardware

JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

DGX agent

arXiv:2605.26636v1 Announce Type: cross Abstract: We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention

hardwarearxiv-cs-ai
27 May 2026
Model Releases

JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors

DGX agent

arXiv:2605.26955v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural c

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

L2Rec: Towards Dual-View Understanding of LLMs for Personalized Recommendation

DGX agent

arXiv:2605.26717v1 Announce Type: cross Abstract: Adapting large language models (LLMs) for personalized recommendation requires aligning their general-purpose capabilities with user-specific preferen

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval

DGX agent

arXiv:2510.13217v2 Announce Type: replace-cross Abstract: Search systems are increasingly used for reasoning-intensive queries, where what makes a document relevant requires understanding or reasoning

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

MemFail: Stress-Testing Failure Modes of LLM Memory Systems

DGX agent

arXiv:2605.26667v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empiric

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

DGX agent

arXiv:2605.27235v1 Announce Type: new Abstract: Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, an

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Neural Autoregressive Control Variates for the Quantum Monte Carlo Sign Problem

DGX agent

arXiv:2605.26814v1 Announce Type: cross Abstract: We train a pair of autoregressive models to construct zero-mean control variates to mitigate the sign problem in quantum Monte Carlo simulations. The

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection

DGX agent

arXiv:2508.01253v2 Announce Type: replace Abstract: Existing studies typically investigate domain shift and category shift as independent problems, however, in real-world scenarios, the two types of s

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs

DGX agent

arXiv:2510.05864v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs

model-releasesarxiv-cs-cl
27 May 2026
Safety

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

DGX agent

arXiv:2605.26526v1 Announce Type: new Abstract: Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an ass

safetyarxiv-cs-lg
27 May 2026
Model Releases

Periodic Topological Deep Learning for Polymer Design and Discovery

DGX agent

arXiv:2605.26833v1 Announce Type: cross Abstract: Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging.

model-releasesarxiv-cs-ai
27 May 2026
Local Ai

Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study

DGX agent

arXiv:2605.26870v1 Announce Type: cross Abstract: Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens wh

local-aiarxiv-cs-ai
27 May 2026
Model Releases

Pretrained Approximators for Low-Thrust Trajectory Cost and Reachability

DGX agent

arXiv:2605.26790v1 Announce Type: new Abstract: Low-thrust trajectory design relies heavily on repeated evaluations of fuel consumption and transfer feasibility, which require expensive optimal contro

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Probing Cultural Awareness in LLMs: A Case Study of Cross-Culture Aesthetic Stylistics

DGX agent

arXiv:2605.27296v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in diverse cultural contexts, yet their ability to master aesthetic stylistics, i.e., the strateg

model-releasesarxiv-cs-cl
27 May 2026
Agents

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

DGX agent

arXiv:2605.27068v1 Announce Type: cross Abstract: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM)

agentsarxiv-cs-ai
27 May 2026
Model Releases

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

DGX agent

arXiv:2605.27134v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking,

model-releasesarxiv-cs-ai
27 May 2026
Hardware

SIA: Self Improving AI with Harness & Weight Updates

DGX agent

arXiv:2605.27276v1 Announce Type: new Abstract: Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The l

hardwarearxiv-cs-ai
27 May 2026
Model Releases

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution

DGX agent

arXiv:2603.01327v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-leve

model-releasesarxiv-cs-cl
27 May 2026
Safety

Unique Lives, Shared World: Learning from Single-Life Videos

DGX agent

arXiv:2512.04085v2 Announce Type: replace Abstract: We introduce the 'single-life' learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual

safetyarxiv-cs-cv
27 May 2026
Model Releases

Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization

DGX agent

arXiv:2605.26433v1 Announce Type: new Abstract: Large language model (LLM) summarization systems may pass compact vector representations of private inputs to downstream retrieval, monitoring, audit, o

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Warp’s big bet on building open source with GPT-5.5

DGX agent

Warp is making a significant investment in developing open source tools and integrations built on GPT-5.5, OpenAI's advanced language model. The initiative aims to leverage GPT-5.5's capabilities to c

model-releasesopenai
27 May 2026
Model Releases

Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis

DGX agent

arXiv:2605.26655v1 Announce Type: new Abstract: Automated prompt optimization methods (e.g., DSpy, TextGrad) can substantially improve the performance of large language model (LLM), however, their gen

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

DGX agent

arXiv:2605.23977v1 Announce Type: new Abstract: This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS,

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AI Content Moderation in Therapy Conversations

DGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

DGX agent

arXiv:2602.22769v3 Announce Type: replace Abstract: Large Language Models (LLMs) are deployed as autonomous agents in increasingly complex applications, where enabling long-horizon memory is critical

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AnnotateMissense: a genome-wide annotation and benchmarking framework for missense pathogenicity prediction

DGX agent

arXiv:2605.24520v1 Announce Type: cross Abstract: Missense variant interpretation remains challenging because pathogenicity depends on heterogeneous evidence from population frequency, evolutionary co

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

DGX agent

arXiv:2605.25626v1 Announce Type: new Abstract: Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its infor

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

DGX agent

arXiv:2605.25920v1 Announce Type: cross Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint

model-releasesarxiv-cs-ai
26 May 2026
Safety

Causal methods for LLM development and evaluation

DGX agent

arXiv:2605.25998v1 Announce Type: new Abstract: Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and

safetyarxiv-cs-lg
26 May 2026
Model Releases

Code2UML: Agentic LLMs with context engineering for scalable software visualization

DGX agent

arXiv:2605.24453v1 Announce Type: cross Abstract: Large Language Model (LLM)-based code analysis tools are adopted to automate software documentation tasks. However, the scalability of these approache

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Cross-Domain Energy-Guided Diffusion Generation for Off-Dynamics Reinforcement Learning

DGX agent

arXiv:2605.24810v1 Announce Type: cross Abstract: Off-dynamics offline reinforcement learning seeks to learn a target-domain policy from a large source dataset and a limited target dataset under misma

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval

DGX agent

arXiv:2605.24454v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance in the legal domain, demonstrating notable potential in Legal Question Answering (LQA). Howev

model-releasesarxiv-cs-cl
26 May 2026
Research

Deep Learning-Enabled Prediction of Geoeffective CMEs Using SOHO and SDO Observations

DGX agent

arXiv:2605.24748v1 Announce Type: cross Abstract: Understanding and forecasting the geoeffectiveness of a coronal mass ejection (CME) is crucial for protecting infrastructure in the near-Earth space e

researcharxiv-cs-lg
26 May 2026
Model Releases

DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

DGX agent

arXiv:2605.26087v1 Announce Type: cross Abstract: Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of establis

model-releasesarxiv-cs-lg
26 May 2026
Research

Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT

DGX agent

arXiv:2605.25924v1 Announce Type: new Abstract: Recent automated essay scoring (AES) studies increasingly use pretrained transformer models, but these models are usually pretrained on general-domain E

researcharxiv-cs-cl
26 May 2026
Model Releases

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

DGX agent

arXiv:2605.23954v1 Announce Type: cross Abstract: Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing

model-releasesarxiv-cs-ai
26 May 2026
← Previous
1…534535536537538…1381
Next →