AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,477 results
24 Apr 2026

Finding Meaning in Embeddings: Concept Separation Curves

ResearchDGX agent

arXiv:2604.21555v1 Announce Type: new Abstract: Sentence embedding techniques aim to encode key concepts of a sentence's meaning in a vector space. However, the majority of evaluation approaches for s

Frequency-Forcing: From Scaling-as-Time to Soft Frequency Guidance

ResearchDGX agent

arXiv:2604.20902v1 Announce Type: cross Abstract: While standard flow-matching models transport noise to data uniformly, incorporating an explicit generation order - specifically, establishing coarse,

From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media

Model ReleasesDGX agent

arXiv:2604.21786v1 Announce Type: new Abstract: Social media platforms have become primary arenas for climate communication, generating millions of images and posts that - if systematically analysed -

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning

Model ReleasesDGX agent

arXiv:2604.21495v1 Announce Type: cross Abstract: Numerical reasoning over expert-domain tables often exhibits high in-domain accuracy but limited robustness to domain shift. Models trained with super

Geometric Characterisation and Structured Trajectory Surrogates for Clinical Dataset Condensation

Model ReleasesDGX agent

arXiv:2604.21638v1 Announce Type: new Abstract: Dataset condensation constructs compact synthetic datasets that retain the training utility of large real-world datasets, enabling efficient model devel

GPT-5.5 is now available in Devin as an Agent Preview! GPT-5.5 has set a new bar for what's possible with Devin. It runs longer and more aut…

Model ReleasesDGX agent

GPT-5.5 is now available in Devin as an Agent Preview! GPT-5.5 has set a new bar for what's possible with Devin. It runs longer and more autonomously than any GPT model we've tested, surfacing bugs no

Grok Voice is used by @Starlink

ApplicationsDGX agent

Grok Voice is used by @Starlink Introducing Grok Voice Think Fast 1.0 A state-of-the-art voice model built for complex, multi-step workflows with snappy responses and high accuracy. It takes the top s

HyperAdapt: Simple High-Rank Adaptation

Model ReleasesDGX agent

arXiv:2509.18629v3 Announce Type: replace-cross Abstract: Foundation models excel across diverse tasks, but adapting them to specialized applications often requires fine-tuning, an approach that is me

Ideological Bias in LLMs' Economic Causal Reasoning

Model ReleasesDGX agent

arXiv:2604.21334v1 Announce Type: new Abstract: Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly used in polic

Language as a Latent Variable for Reasoning Optimization

Model ReleasesDGX agent

arXiv:2604.21593v1 Announce Type: new Abstract: As LLMs reduce English-centric bias, a surprising trend emerges: non-English responses sometimes outperform English on reasoning tasks. We hypothesize t

Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery

Model ReleasesDGX agent

arXiv:2604.21102v1 Announce Type: cross Abstract: We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLM

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations

Model ReleasesDGX agent

arXiv:2509.25868v3 Announce Type: replace Abstract: The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And C

Slot Machines: How LLMs Keep Track of Multiple Entities

ResearchDGX agent

arXiv:2604.21139v1 Announce Type: new Abstract: Language models must bind entities to the attributes they possess and maintain several such binding relationships within a context. We study how multipl

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

Model ReleasesDGX agent

arXiv:2604.21700v1 Announce Type: cross Abstract: The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studie

Strategic Scaling of Test-Time Compute: A Bandit Learning Approach

TutorialsDGX agent

arXiv:2506.12721v2 Announce Type: replace Abstract: Scaling test-time compute has emerged as an effective strategy for improving the performance of large language models. However, existing methods typ

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning

Model ReleasesDGX agent

arXiv:2604.21327v1 Announce Type: cross Abstract: Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization sig

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs

Model ReleasesDGX agent

arXiv:2604.21911v1 Announce Type: cross Abstract: Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs

23 Apr 2026

Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2604.19974v1 Announce Type: cross Abstract: Large language models can be uncertain yet correct, or confident yet wrong, raising the question of whether their output-level uncertainty and their a

Best Policy Learning from Trajectory Preference Feedback

SafetyDGX agent

arXiv:2501.18873v4 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned rew

Beyond ZOH: Advanced Discretization Strategies for Vision Mamba

ResearchDGX agent

arXiv:2604.20606v1 Announce Type: cross Abstract: Vision Mamba, as a state space model (SSM), employs a zero-order hold (ZOH) discretization, which assumes that input signals remain constant between s

CRAFT: Training-Free Cascaded Retrieval for Tabular QA

Model ReleasesDGX agent

arXiv:2505.14984v2 Announce Type: replace Abstract: Open-Domain Table Question Answering (TQA) involves retrieving relevant tables from a large corpus to answer natural language queries. Traditional d

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data

AgentsDGX agent

arXiv:2604.19859v1 Announce Type: cross Abstract: Edge-scale deep research agents based on small language models are attractive for real-world deployment due to their advantages in cost, latency, and

Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing

Model ReleasesDGX agent

arXiv:2604.20429v1 Announce Type: new Abstract: Remote sensing (RS) image-text retrieval plays a critical role in understanding massive RS imagery. However, the dense multi-object distribution and com

Knowledge Capsules: Structured Nonparametric Memory Units for LLMs

Model ReleasesDGX agent

arXiv:2604.20487v1 Announce Type: cross Abstract: Large language models (LLMs) encode knowledge in parametric weights, making it costly to update or extend without retraining. Retrieval-augmented gene

Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering

Model ReleasesDGX agent

arXiv:2604.16756v2 Announce Type: replace-cross Abstract: Prompt-induced cognitive biases are changes in a general-purpose AI (GPAI) system's decisions caused solely by biased wording in the input (e.

🚨 OpenAI just launched GPT-5.5. The OpenAI team was nice enough to give me early access over the last several weeks, and I just want to fla…

Model ReleasesDGX agent

🚨 OpenAI just launched GPT-5.5. The OpenAI team was nice enough to give me early access over the last several weeks, and I just want to flag: there is a certain class of models (one that we’re hitting

Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL

Model ReleasesDGX agent

arXiv:2604.20835v1 Announce Type: new Abstract: Modern language models demonstrate impressive coding capabilities in common programming languages (PLs), such as C++ and Python, but their performance i

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

ResearchDGX agent

arXiv:2604.20696v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they sti

Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

TutorialsDGX agent

arXiv:2604.20730v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. Howev

Semantic-Fast-SAM: Efficient Semantic Segmenter

ResearchDGX agent

arXiv:2604.20169v1 Announce Type: new Abstract: We propose Semantic-Fast-SAM (SFS), a semantic segmentation framework that combines the Fast Segment Anything model with a semantic labeling pipeline to

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation

Model ReleasesDGX agent

arXiv:2604.20842v1 Announce Type: cross Abstract: Paralinguistic cues are essential for natural human-computer interaction, yet their evaluation in Large Audio-Language Models (LALMs) remains limited

SweRank: Software Issue Localization with Code Ranking

Model ReleasesDGX agent

arXiv:2505.07849v2 Announce Type: replace-cross Abstract: Software issue localization, the task of identifying the precise code locations (files, classes, or functions) relevant to a natural language

Think fast!

ApplicationsDGX agent

Think fast! Introducing Grok Voice Think Fast 1.0 A state-of-the-art voice model built for complex, multi-step workflows with snappy responses and high accuracy. It takes the top spot on the Tau Voice

Transformers Can Learn Connectivity in Some Graphs but Not Others

TutorialsDGX agent

arXiv:2509.22343v2 Announce Type: replace-cross Abstract: Reasoning capability is essential to ensure the factual correctness of the responses of transformer-based Large Language Models (LLMs), and ro

22 Apr 2026

a bunch here where I’m saying ok Garry’s kinda right?! 👀…in some ways :) we’re making this loop much easier to close out of the box soon If…

Model ReleasesDGX agent

a bunch here where I’m saying ok Garry’s kinda right?! 👀…in some ways :) we’re making this loop much easier to close out of the box soon If more people get into evals & traces to ground self-improving

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison

Model ReleasesDGX agent

arXiv:2603.13779v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive success in natural visual understanding, yet they consistently underperform

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

SafetyDGX agent

arXiv:2604.18789v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an im

BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design

ResearchDGX agent

arXiv:2508.21184v3 Announce Type: replace-cross Abstract: We propose a general-purpose approach for improving the ability of large language models (LLMs) to intelligently and adaptively gather informa

Benchmarking Misuse Mitigation Against Covert Adversaries

SafetyDGX agent

arXiv:2506.06414v2 Announce Type: replace-cross Abstract: Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing sa

Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks

Model ReleasesDGX agent

arXiv:2512.22673v3 Announce Type: replace Abstract: Travel planning is a natural real-world task to test large language models' (LLMs) planning and tool-use abilities. Although prior work has studied

Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs

SafetyDGX agent

arXiv:2601.15755v3 Announce Type: replace Abstract: Large language models are increasingly used to represent human opinions, values, or beliefs, and their steerability towards these ideals is an activ

Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning

Model ReleasesDGX agent

arXiv:2604.18715v1 Announce Type: cross Abstract: Earth observation foundation models encode land surface information into dense embedding vectors, yet the geometric structure of these representations

Day 1 at Google Cloud Next ‘26 recap

Model ReleasesDGX agent

Last year at Google Cloud Next ‘25, we asked you to imagine a new future for AI. At Next ‘26, the question before you is how do you move AI into production across your entire enterprise? According to

Deep Supervised Contrastive Learning of Pitch Contours for Robust Pitch Accent Classification in Seoul Korean

Model ReleasesDGX agent

arXiv:2604.19477v1 Announce Type: cross Abstract: The intonational structure of Seoul Korean has been defined with discrete tonal categories within the Autosegmental-Metrical model of intonational pho

Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning

Model ReleasesDGX agent

arXiv:2604.19459v1 Announce Type: new Abstract: Formal verification guarantees proof validity but not formalization faithfulness. For natural-language logical reasoning, where models construct axiom s

Environmental Sound Deepfake Detection Using Deep-Learning Framework

Model ReleasesDGX agent

arXiv:2604.19652v1 Announce Type: cross Abstract: In this paper, we propose a deep-learning framework for environmental sound deepfake detection (ESDD) -- the task of identifying whether the sound sce

Evaluation-driven Scaling for Scientific Discovery

Local AiDGX agent

arXiv:2604.19341v1 Announce Type: cross Abstract: Language models are increasingly used in scientific discovery to generate hypotheses, propose candidate solutions, implement systems, and iteratively

Fine-tuning DeepSeek-OCR-2 for Molecular Structure Recognition

Model ReleasesDGX agent

arXiv:2604.03476v2 Announce Type: replace-cross Abstract: Optical Chemical Structure Recognition (OCSR) is critical for converting 2D molecular diagrams from printed literature into machine-readable f

LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit

SafetyDGX agent

arXiv:2604.19117v1 Announce Type: new Abstract: When a language model agrees with a user's false belief, is it failing to detect the error, or noticing and agreeing anyway? We show the latter. Across

LSTM-MAS: A Long Short-Term Memory Inspired Multi-Agent System for Long-Context Understanding

Model ReleasesDGX agent

arXiv:2601.11913v2 Announce Type: replace-cross Abstract: Effectively processing long contexts remains a fundamental yet unsolved challenge for large language models (LLMs). Existing single-LLM-based

MapPFN: Learning Causal Perturbation Maps in Context

TutorialsDGX agent

arXiv:2601.21092v2 Announce Type: replace Abstract: Planning effective interventions in biological systems requires treatment-effect models that adapt to unseen biological contexts by identifying thei

Mechanistic Anomaly Detection via Functional Attribution

Model ReleasesDGX agent

arXiv:2604.18970v1 Announce Type: new Abstract: We can often verify the correctness of neural network outputs using ground truth labels, but we cannot reliably determine whether the output was produce

Multi-Domain Learning with Global Expert Mapping

Model ReleasesDGX agent

arXiv:2604.18842v1 Announce Type: new Abstract: Human perception generalizes well across different domains, but most vision models struggle beyond their training data. This gap motivates multi-dataset

Optimal Routing for Federated Learning over Dynamic Satellite Networks: Tractable or Not?

Local AiDGX agent

arXiv:2604.19399v1 Announce Type: new Abstract: Federated learning (FL) is a key paradigm for distributed model learning across decentralized data sources. Communication in each FL round typically con

PriorGuide: Test-Time Prior Adaptation for Simulation-Based Inference

Model ReleasesDGX agent

arXiv:2510.13763v2 Announce Type: replace-cross Abstract: Amortized simulator-based inference offers a powerful framework for tackling Bayesian inference in computational fields such as engineering or

Probing for Reading Times

SafetyDGX agent

arXiv:2604.18712v1 Announce Type: new Abstract: Probing has shown that language model representations encode rich linguistic information, but it remains unclear whether they also capture cognitive sig

Safe Continual Reinforcement Learning in Non-stationary Environments

Model ReleasesDGX agent

arXiv:2604.19737v1 Announce Type: new Abstract: Reinforcement learning (RL) offers a compelling data-driven paradigm for synthesizing controllers for complex systems when accurate physical models are

SimDiff: Depth Pruning via Similarity and Difference

ResearchDGX agent

arXiv:2604.19520v1 Announce Type: new Abstract: Depth pruning improves the deployment efficiency of large language models (LLMs) by identifying and removing redundant layers. A widely accepted standar

Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

Model ReleasesDGX agent

arXiv:2604.19548v1 Announce Type: cross Abstract: Large Language Model agents have rapidly evolved from static text generators into dynamic systems capable of executing complex autonomous workflows. T

Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark

Model ReleasesDGX agent

arXiv:2511.01233v3 Announce Type: replace Abstract: We review human evaluation practices in automatic, speech-driven 3D gesture generation and find a lack of standardisation and frequent use of flawed

← Previous
1…319320321322323…1042
Next →