AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
12 May 2026

Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models

Model ReleasesDGX agent

arXiv:2605.09496v1 Announce Type: new Abstract: Large language models represent the same reasoning in vastly different surface forms -- English prose, Python code, mathematical notation -- yet whether

Beyond Local Edits: Embedding-Virtualized Knowledge for Broader Evaluation and Preservation of Model Editing

Model ReleasesDGX agent

arXiv:2602.01977v2 Announce Type: replace Abstract: Knowledge editing methods for large language models are commonly evaluated using predefined benchmarks that assess edited facts together with a limi

Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.10901v1 Announce Type: new Abstract: Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal

Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10034v1 Announce Type: new Abstract: Recent Autonomous Driving (AD) works such as GigaFlow and PufferDrive have unlocked Reinforcement Learning (RL) at scale as a training strategy for driv

Beyond source code: The files AI coding agents trust — and attackers exploit

Model ReleasesDGX agent

As AI coding agents become deeply embedded in developer workflows, defenders must evolve their definition of malicious files and rethink how to protect against them. Autonomous AI agents operate acros

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

Model ReleasesDGX agent

arXiv:2605.08761v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles,

Beyond the False Trade-off: Adaptive EWC for Stealthy and Generalizable T2I Backdoors

Model ReleasesDGX agent

arXiv:2605.08280v1 Announce Type: cross Abstract: Preserving model fidelity is essential for stealthy text-to-image (T2I) backdoor attacks. Existing methods such as Learning without Forgetting (LwF) r

Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation

Model ReleasesDGX agent

arXiv:2502.08943v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural l

Beyond Toy Benchmarks: A Systematic Evaluation of OOD Detection Methods For Plant Pathology Classification

Model ReleasesDGX agent

arXiv:2605.08618v1 Announce Type: new Abstract: Out-of-distribution (OOD) detection is essential for reliable deployment of deep learning systems, yet the majority of existing methods are evaluated on

BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization

Model ReleasesDGX agent

arXiv:2605.10345v1 Announce Type: new Abstract: Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence

Model ReleasesDGX agent

arXiv:2605.09041v1 Announce Type: new Abstract: Bias audits of large language models now operate within governance frameworks such as the EU AI Act, making benchmark reliability a security concern in

Bilinear autoencoders find interpretable manifolds

Model ReleasesDGX agent

arXiv:2605.08891v1 Announce Type: new Abstract: Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span

BoostLLM: Boosting-inspired LLM Fine-tuning for Few-shot Tabular Classification

Model ReleasesDGX agent

arXiv:2605.06117v2 Announce Type: replace Abstract: Large language models (LLMs) have recently been adapted to tabular prediction by serializing structured features into natural language, but their pe

Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

Model ReleasesDGX agent

arXiv:2605.08271v1 Announce Type: cross Abstract: Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For

Bridging Spectral Operator Learning and U-Net Hierarchies: SpectraNet for Stable Autoregressive PDE Surrogates

Model ReleasesDGX agent

arXiv:2605.09096v1 Announce Type: new Abstract: Neural operators for time-dependent PDEs face a structural tension: spectral architectures (FNO and descendants) inherit exponential rollout-error growt

Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models

Model ReleasesDGX agent

arXiv:2605.08404v1 Announce Type: cross Abstract: This work investigates the use of large language models (LLMs) for tasks in smart cities. The core idea is to leverage remote sensing imagery to chara

bViT: Investigating Single-Block Recurrence in Vision Transformers for Image Recognition

Model ReleasesDGX agent

arXiv:2605.10661v1 Announce Type: cross Abstract: Vision Transformers (ViTs) are built by stacking independently parameterized blocks, but it remains unclear how much of this depth requires layer spec

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks

Model ReleasesDGX agent

arXiv:2605.09611v1 Announce Type: new Abstract: This preprint presents an empirical analysis of byte-exact chunk-level deduplication in Retrieval-Augmented Generation (RAG) pipelines. We measure conte

C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10744v1 Announce Type: new Abstract: Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

Model ReleasesDGX agent

arXiv:2605.10873v1 Announce Type: cross Abstract: Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure because existin

CALYREX: Cross-Attention LaYeR EXtended Transformers for System Prompt Anchoring

Model ReleasesDGX agent

arXiv:2605.09737v1 Announce Type: new Abstract: Modern large language models (LLMs) rely on system prompts to establish behavioral constraints and safety rules. Standard causal self-attention treats p

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.10448v1 Announce Type: new Abstract: Interactive agent benchmarks map an agent run to a binary outcome through outcome checks. When these checks rely on surface level signals or fail to cap

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

Model ReleasesDGX agent

arXiv:2601.12369v3 Announce Type: replace Abstract: Deep Research Agents increasingly automate survey generation, yet whether they match human experts at retrieving essential papers and organizing the

Can LLMs Predict Polymer Physics Just by Reading Synthesis and Processing Prose?

Model ReleasesDGX agent

arXiv:2605.08255v1 Announce Type: cross Abstract: Can large language models predict physical and mechanical polymer properties simply by reading unstructured scientific prose? Polymer performance is r

Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology

Model ReleasesDGX agent

arXiv:2512.06949v3 Announce Type: replace Abstract: Histopathology image segmentation is essential for delineating tissue structures in skin cancer diagnostics, but modeling spatial context and inter-

Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness

Model ReleasesDGX agent

arXiv:2605.09634v1 Announce Type: new Abstract: LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability ac

Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts

Model ReleasesDGX agent

arXiv:2503.05066v5 Announce Type: replace-cross Abstract: The Mixture of Experts (MoE) is an effective architecture for scaling large language models by leveraging sparse expert activation to balance

CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.10903v1 Announce Type: new Abstract: This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adapta

Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure

Model ReleasesDGX agent

arXiv:2605.08740v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) decompose transformer residual streams into interpretable feature dictionaries, yet the relationship between SAE width and

Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations

Model ReleasesDGX agent

arXiv:2605.08590v1 Announce Type: cross Abstract: LLMs are increasingly used to explain personal sensing data, translating traces of activity and mood into natural-language accounts of why an anomalou

Cavity-Enhanced Collective Quantum Processing with Polarization-Encoded Qubits

Model ReleasesDGX agent

arXiv:2605.10473v1 Announce Type: cross Abstract: We introduce a cavity-enhanced optical architecture for collective quantum processing in which logical qubits are encoded in the polarization subspace

CERSA: Cumulative Energy-Retaining Subspace Adaptation for Memory-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.08174v1 Announce Type: cross Abstract: To mitigate the memory constraints associated with fine-tuning large pre-trained models, existing parameter-efficient fine-tuning (PEFT) methods, such

CHAINTRIX: A multi-pipeline LLM-augmented framework for automated smart-contract security auditing

Model ReleasesDGX agent

arXiv:2605.09350v1 Announce Type: new Abstract: Smart-contract exploits have caused billions of USD in cumulative losses, yet audits remain expensive and slow. Automated tools have emerged to close th

ChaosNetBench: Benchmarking Spatio-Temporal Graph Neural Networks on Chaotic Lattice Dynamics

Model ReleasesDGX agent

arXiv:2605.09676v1 Announce Type: cross Abstract: Spatio-temporal graph neural networks (STGNNs) are widely used for short-term forecasting in dynamic physical systems such as traffic and weather. How

Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus

Model ReleasesDGX agent

arXiv:2605.09092v1 Announce Type: new Abstract: This study addresses automatic transliteration from Tajik (Cyrillic script) to Persian (Perso-Arabic script). We present a curated, lexicographically ve

ChartDiff: A Large-Scale Benchmark for Comprehending Pairs of Charts

Model ReleasesDGX agent

arXiv:2603.28902v2 Announce Type: replace Abstract: Charts are central to analytical reasoning, yet existing benchmarks for chart understanding focus almost exclusively on single-chart interpretation

ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour

Model ReleasesDGX agent

arXiv:2506.12090v2 Announce Type: replace Abstract: This paper introduces ChatbotManip, a novel dataset for studying manipulation in Chatbots. It contains simulated generated conversations between a c

Chinese Cyberbullying Detection: Dataset, Method, and Validation

Model ReleasesDGX agent

arXiv:2505.20654v2 Announce Type: replace-cross Abstract: Existing cyberbullying detection benchmarks were organized by the polarity of speech, such as 'offensive' and 'non-offensive', which were esse

CIVeX: Causal Intervention Verification for Language Agents

Model ReleasesDGX agent

arXiv:2605.09168v1 Announce Type: new Abstract: A valid tool call is not necessarily a valid intervention. Tool-using language agents are guarded by schema validators, policy filters, provenance check

Classification-Head Bias in Class-Level Machine Unlearning: Diagnosis, Mitigation, and Evaluation

Model ReleasesDGX agent

arXiv:2605.08730v1 Announce Type: new Abstract: Class-level machine unlearning aims to remove the influence of specified classes while preserving model utility on retained classes. Existing methods ar

🤩🤯🤩 Claude Code (still not AGI but biggest advance since GPT-4) is the most neurosymbolic thing I have ever seen in my life. 53 symbolic …

Model ReleasesDGX agent

🤩🤯🤩 Claude Code (still not AGI but biggest advance since GPT-4) is the most neurosymbolic thing I have ever seen in my life. 53 symbolic tools, 500,000 lines of symbolic code, combined with a state-of

Claude Opus 4.7 (fast mode) is now available in Windsurf! Full Claude Opus 4.7 intelligence ~2.5x higher output speeds We can't wait to see …

Model ReleasesDGX agent

Windsurf has released Claude Opus 4.7 in fast mode, offering the full intelligence of Claude Opus 4.7 with approximately 2.5x faster output speeds. This integration allows users of Windsurf to access

CLEF: EEG Foundation Model for Learning Clinical Semantics

Model ReleasesDGX agent

arXiv:2605.10817v1 Announce Type: new Abstract: Clinical EEG interpretation requires reasoning over full EEG sessions and integrating signal patterns with clinical context. Existing EEG foundation mod

Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories

Model ReleasesDGX agent

arXiv:2605.10840v1 Announce Type: cross Abstract: We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories. JEPA architec

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics

Model ReleasesDGX agent

arXiv:2605.09584v1 Announce Type: cross Abstract: Inpatient clinical reasoning is a sequential decision under partial observability: the clinician sees the admission so far and must choose the next ac

CMKL: Modality-Aware Continual Learning for Evolving Biomedical Knowledge Graphs

Model ReleasesDGX agent

arXiv:2605.10510v1 Announce Type: cross Abstract: Biomedical knowledge graphs are increasingly large, dynamic, and multimodal, driven by rapid advances in biotechnology such as high-throughput sequenc

CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis

Model ReleasesDGX agent

arXiv:2510.11233v3 Announce Type: replace Abstract: Depression is a pressing global public health issue, yet publicly available Chinese-language resources for depression risk detection remain scarce a

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents

Model ReleasesDGX agent

arXiv:2605.09675v1 Announce Type: new Abstract: Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tra

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization

Model ReleasesDGX agent

arXiv:2605.08873v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has emerged as a powerful algorithm for improving the reasoning capabilities of language models, but often fai

Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

Model ReleasesDGX agent

arXiv:2605.10528v1 Announce Type: cross Abstract: We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physi

ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2601.16836v3 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to a

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization

Model ReleasesDGX agent

arXiv:2605.08802v1 Announce Type: new Abstract: Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perfo

Combining Mechanical and Agentic Specification Inference for Move

Model ReleasesDGX agent

arXiv:2605.10005v1 Announce Type: cross Abstract: In this paper, we describe early work on a specification inference tool for the Move Prover that combines a weakest-precondition (WP) analysis over Mo

ComfyUI Support for HiDream-01-Image Released

Model ReleasesDGX agent

ComfyUI support for HiDream-O1-Image enables local image generation with text prompts and optional reference images, featuring various precision options (BF16/FP16/FP32/FP8) and integration with atten

Communication Dynamics Neural Networks: FFT-Diagonalized Layers for Improved Hessian Conditioning at Reduced Parameter Count

Model ReleasesDGX agent

arXiv:2605.08171v1 Announce Type: cross Abstract: Background and motivation. The Communication Dynamics (CD) framework, introduced in two earlier papers for atomic-energy prediction and field-induced

Compander-Aligned Query Geometry for Quantized Zeroth-Order Optimization

Model ReleasesDGX agent

arXiv:2605.10673v1 Announce Type: new Abstract: Low-bit forward evaluation is an attractive route to memory-efficient zeroth-order (ZO) adaptation: the optimizer needs only scalar losses, and the mode

Complex-Valued Phase-Coherent Transformer

Model ReleasesDGX agent

arXiv:2605.10123v1 Announce Type: new Abstract: Complex-valued Transformers have largely inherited softmax attention from real-valued architectures. However, row-normalised token competition is not ne

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

Model ReleasesDGX agent

arXiv:2605.10787v1 Announce Type: new Abstract: Current LLM agents are proficient at calling isolated APIs but struggle with the 'last mile' of commercial software automation. In real-world scenarios,

Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation

Model ReleasesDGX agent

arXiv:2605.08810v1 Announce Type: cross Abstract: We propose Compressed Video Aggregator (CVA), a lightweight micro-video recommendation module that decouples video information from preference learnin

Computer Use at the Edge of the Statistical Precipice

Model ReleasesDGX agent

arXiv:2605.08261v1 Announce Type: cross Abstract: Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically add

← Previous
1…260261262263264…377
Next →