AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
91,020Total entries
1Added by human
91,019Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,766 results
Safety

How Value Induction Reshapes LLM Behaviour

DGX agent

arXiv:2605.07925v1 Announce Type: new Abstract: Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and em

safetyarxiv-cs-cl
11 May 2026
Model Releases

Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.05957v2 Announce Type: replace Abstract: LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply r

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System

DGX agent

arXiv:2605.05949v2 Announce Type: replace Abstract: Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text

DGX agent

arXiv:2605.06903v1 Announce Type: cross Abstract: Large language models are now embedded in everyday writing workflows, making reliable AI-generated text detection important for academic integrity, co

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

On the Invariance and Generality of Neural Scaling Laws

DGX agent

arXiv:2605.07546v1 Announce Type: new Abstract: Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocatio

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Pretraining Induces a Reusable Spectral Basis for Downstream Task Adaptation

DGX agent

arXiv:2605.07302v1 Announce Type: new Abstract: Finetuning pretrained models occurs in a low-dimensional subspace of the full parameter space. Prior work has focused on characterizing this optimizatio

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

DGX agent

arXiv:2605.07141v1 Announce Type: cross Abstract: Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large lang

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

DGX agent

arXiv:2605.05995v2 Announce Type: replace-cross Abstract: The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constrain

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation

DGX agent

arXiv:2603.05117v3 Announce Type: replace Abstract: Imitation Learning (IL) enables robots to acquire manipulation skills from expert demonstrations. Diffusion Policy (DP) models multi-modal expert be

model-releasesarxiv-cs-ro
11 May 2026
Research

Synergistic Benefits of Joint Molecule Generation and Property Prediction

DGX agent

arXiv:2504.16559v3 Announce Type: replace Abstract: Modeling the joint distribution of data samples and their properties allows to construct a single model for both data generation and property predic

researcharxiv-cs-lg
11 May 2026
Model Releases

Test-Time Compute Games

DGX agent

arXiv:2601.21839v2 Announce Type: replace-cross Abstract: Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strate

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List

DGX agent

arXiv:2605.07127v1 Announce Type: cross Abstract: Modern large language models (LLMs) can find a needle in a haystack (locating a single relevant fact buried among hundreds of thousands of irrelevant

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

DGX agent

arXiv:2605.06772v1 Announce Type: new Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical questi

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

DGX agent

arXiv:2605.07114v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language model

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Why Self-Inconsistency Arises in GNN Explanations and How to Exploit It

DGX agent

arXiv:2605.07527v1 Announce Type: cross Abstract: Recent work has observed that explanations produced by Self-Interpretable Graph Neural Networks (SI-GNNs) can be self-inconsistent: when the model is

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

ERNIE 5.1 is here 🚀 ERNIE 5.1 significantly reduces pretraining cost while compressing total parameters to ~1/3 and activated parameters to…

DGX agent

ERNIE 5.1 is here 🚀 ERNIE 5.1 significantly reduces pretraining cost while compressing total parameters to ~1/3 and activated parameters to ~1/2 — using only ~6% of the pretraining cost compared to mo

model-releasesjeremy-howard--x
9 May 2026
Model Releases

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)

DGX agent

arXiv:2605.04556v1 Announce Type: cross Abstract: The Massive Sound Embedding Benchmark (MSEB) has emerged as a standard for evaluating the functional breadth of audio models. While initial baselines

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Benchmarking POS Tagging for the Tajik Language: A Comparative Study of Neural Architectures on the TajPersParallel Corpus

DGX agent

arXiv:2605.04576v1 Announce Type: new Abstract: This paper presents the first benchmark for the task of automatic part-of-speech (POS) tagging for the Tajik language. Despite the existence of multilin

model-releasesarxiv-cs-cl
7 May 2026
Hardware

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs

DGX agent

arXiv:2605.04357v1 Announce Type: cross Abstract: The usage of large language models (LLMs) has grown increasingly fragmented, with no single model dominating. Meanwhile, cloud providers offer a wide

hardwarearxiv-cs-cl
7 May 2026
Model Releases

Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

DGX agent

arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environm

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression

DGX agent

arXiv:2605.04084v1 Announce Type: new Abstract: Compressing large language models (LLMs) for deployment on commodity GPUs remains challenging: conventional scalar quantization is limited to fixed bit-

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

FlatASCEND: Autoregressive Clinical Sequence Generation with Continuous Time Prediction and Association-Based Pharmacological Testing

DGX agent

arXiv:2605.04071v1 Announce Type: new Abstract: Autoregressive models can predict clinical events, but generating patient-conditioned multi-step trajectories that respond to intervention tokens and te

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs

DGX agent

arXiv:2605.04065v1 Announce Type: new Abstract: Unsupervised reinforcement learning (RL) has emerged as a promising paradigm for enabling self-improvement in large language models (LLMs). However, exi

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages

DGX agent

arXiv:2605.04208v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive multilingual capabilities for well-resourced languages, yet their performance on low-resource

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval

DGX agent

arXiv:2605.03361v2 Announce Type: new Abstract: As multimodal content continues to expand at a rapid pace, audio retrieval has emerged as a key enabling technology for media search, content organizati

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

Tree-Conditioned Edit Flows for Ancestral Sequence Reconstruction

DGX agent

arXiv:2605.04119v1 Announce Type: cross Abstract: Ancestral sequence reconstruction (ASR) aims to infer extinct protein sequences at internal nodes of a phylogenetic tree. Classical ASR methods are ty

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks

DGX agent

arXiv:2605.03759v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs) offer powerful capabilities, they pose privacy risks by unintentionally memorizing sensitive personal informa

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.03310v1 Announce Type: cross Abstract: Multi-agent LLM systems fail in production at rates between 41% and 87%, mostly due to coordination defects rather than base-model capability. Existin

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

Feature-Augmented Transformers for Robust AI-Text Detection Across Domains and Generators

DGX agent

arXiv:2605.03969v1 Announce Type: new Abstract: AI-generated text is nowadays produced at scale across domains and heterogeneous generation pipelines, making robustness to distribution shift a central

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Hello again, everyone! Our latest Qwopus3.6-35B-A3B-v1 is now live, and it is once again breathtaking! Full HF space benchmark showcase and …

DGX agent

Hello again, everyone! Our latest Qwopus3.6-35B-A3B-v1 is now live, and it is once again breathtaking! Full HF space benchmark showcase and write-up is in the comments, so you can make conclusions for

model-releasesclem-delangue--x
6 May 2026
Model Releases

PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

DGX agent

arXiv:2605.03571v1 Announce Type: new Abstract: Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising applicati

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Raising the Ceiling: Better Empirical Fixation Densities for Saliency Benchmarking

DGX agent

arXiv:2605.03885v1 Announce Type: new Abstract: Empirical fixation densities, spatial distributions estimated from human eye-tracking data, are foundational to saliency benchmarking. They directly sha

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

ReCode: Reinforcing Code Generation with Reasoning-Process Rewards

DGX agent

arXiv:2508.05170v3 Announce Type: replace-cross Abstract: In practice, rigorous reasoning is often a key driver of correct code, while Reinforcement Learning (RL) for code generation often neglects op

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning

DGX agent

arXiv:2605.03189v1 Announce Type: new Abstract: Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While sev

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning

DGX agent

arXiv:2605.03229v1 Announce Type: new Abstract: Adapting a pretrained language model to a new task often hurts the general capabilities it already had, a problem known as catastrophic forgetting. Spar

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

DGX agent

arXiv:2605.03792v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar e

model-releasesarxiv-cs-cl
6 May 2026
Research

TsallisPGD: Adaptive Gradient Weighting for Adversarial Attacks on Semantic Segmentation

DGX agent

arXiv:2605.03405v1 Announce Type: new Abstract: Attacking semantic segmentation models is significantly harder than image classification models because an attacker must flip thousands of pixel predict

researcharxiv-cs-cv
6 May 2026
Model Releases

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning

DGX agent

arXiv:2605.00425v1 Announce Type: new Abstract: Reinforcement learning (RL) has significantly advanced the ability of large language model (LLM) agents to interact with environments and solve multi-tu

model-releasesarxiv-cs-ai
5 May 2026
Model Releases

Ai2 releases MolmoAct 2, enhancing robot intelligence in the real world

DGX agent

Seattle-based artificial intelligence research institute Ai2, the Allen Institute for AI, today announced its next-generation open-source foundation artificial intelligence models, aimed at enabling r

model-releasessiliconangle
5 May 2026
Safety

Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation

DGX agent

arXiv:2605.01772v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, enabling robots to perform tasks based on natural l

safetyarxiv-cs-lg
5 May 2026
Model Releases

BIM Information Extraction Through LLM-based Adaptive Exploration

DGX agent

arXiv:2605.01698v1 Announce Type: new Abstract: BIM models provide structured representations of building geometry, semantics, and topology, yet extracting specific information from them remains remar

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget

DGX agent

arXiv:2605.01298v1 Announce Type: cross Abstract: Backdoor attacks threaten the deep learning supply chain by poisoning a small fraction of the training data so that a model behaves normally on clean

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming

DGX agent

arXiv:2605.02647v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety alignment and elicit harmful responses. A growing body of work sh

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features

DGX agent

arXiv:2602.10437v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) decompose language model activations into interpretable features, but existing methods reveal only which features a

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning

DGX agent

arXiv:2509.22075v5 Announce Type: replace Abstract: Post-training compression of large language models (LLMs) often relies on low-rank weight approximations that represent each column of the weight ma

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

DBLP: Phase-Aware Bounded-Loss Transport for Burst-Resilient Distributed ML Training

DGX agent

arXiv:2605.01989v1 Announce Type: new Abstract: Distributed machine learning (ML) training has become a necessity with the prevalence of billion to trillion-parameter-scale models. While prior work ha

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Deep neural networks with Fisher vector encoding for medical image classification

DGX agent

arXiv:2605.01667v1 Announce Type: new Abstract: Orderless encoding methods have shown to improve Convolutional Neural Networks (CNNs) for image classification in the context of limited availability of

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Diet Your LLM: Dimension-wise Global Pruning of LLMs via Merging Task-specific Importance Score

DGX agent

arXiv:2603.23985v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated remarkable capabilities, but their massive scale poses significant challenges for practical deploymen

model-releasesarxiv-cs-lg
5 May 2026
← Previous
1…418419420421422…1371
Next →