AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
12 May 2026

Concordia: Self-Improving Synthetic Tables for Federated LLMs

Model ReleasesDGX agent

arXiv:2605.09855v1 Announce Type: new Abstract: Federated learning (FL) enables training large language models (LLMs) without sharing raw data, but adapting LLMs under strict data isolation and non-II

Confidence-Guided Diffusion Augmentation for Enhanced Bangla Compound Character Recognition

Model ReleasesDGX agent

arXiv:2605.10916v1 Announce Type: cross Abstract: Recognition of handwritten Bangla compound characters remains a challenging problem due to complex character structures, large intra-class variation,

ConFit v3: Improving Resume-Job Matching with LLM-based Re-Ranking

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.09760v1 Announce Type: new Abstract: A reliable resume-job matching system helps a company find suitable candidates from a pool of resumes and helps a job seeker find relevant jobs from a l

ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs

Model ReleasesDGX agent

arXiv:2605.10793v1 Announce Type: new Abstract: Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce

Context-Augmented Code Generation: How Product Context Improves AI Coding Agent Decision Compliance by 49%

Model ReleasesDGX agent

arXiv:2605.08112v1 Announce Type: cross Abstract: AI coding agents powered by large language models can read codebases and produce functional code, but they routinely violate team-specific product dec

Continual Harness: Online Adaptation for Self-Improving Foundation Agents

Model ReleasesDGX agent

arXiv:2605.09998v1 Announce Type: cross Abstract: Coding harnesses such as Claude Code and OpenHands wrap foundation models with tools, memory, and planning, but no equivalent exists for embodied agen

Continuous Latent Contexts Enable Efficient Online Learning in Transformers

Model ReleasesDGX agent

arXiv:2605.09867v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit a strong capacity for in-context learning: Given labeled examples, they can generate good predictions without par

Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit

Model ReleasesDGX agent

arXiv:2605.08352v1 Announce Type: new Abstract: A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. As the number o

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation

Model ReleasesDGX agent

arXiv:2605.08522v1 Announce Type: new Abstract: The evaluation of Large Language Models (LLMs) faces a critical challenge in construct validity, where fragmented benchmarks and ad hoc metrics frequent

CORTEG: Foundation Models Enable Cross-Modality Representation Transfer from Scalp to Intracranial Brain Recordings

Model ReleasesDGX agent

arXiv:2605.10337v1 Announce Type: new Abstract: Intracranial electrocorticography (ECoG) offers high-signal-to-noise access to cortical activity for brain-computer interfaces, yet limited per-patient

Cosine-Gated Adam-Decay: Drop-In Staleness-Aware Outer Optimization for Decoupled DiLoCo

Model ReleasesDGX agent

arXiv:2605.09126v1 Announce Type: new Abstract: Asynchronous DiLoCo systems may receive pseudo-gradients computed several outer rounds earlier, yet the standard Nesterov outer optimizer does not expli

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10426v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms sti

CrackMeBench: Binary Reverse Engineering for Agents

Model ReleasesDGX agent

arXiv:2605.10597v1 Announce Type: cross Abstract: Benchmarks for coding agents increasingly measure source-level software repair, and cybersecurity benchmarks increasingly measure broad capture-the-fl

CREATE: Testing LLMs for Associative Creativity

Model ReleasesDGX agent

arXiv:2603.09970v2 Announce Type: replace Abstract: A key component of creativity is associative reasoning: the ability to draw novel yet meaningful connections between concepts. We introduce CREATE,

Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations

Model ReleasesDGX agent

arXiv:2605.09875v1 Announce Type: new Abstract: Large language models from different families use different hidden dimensions, tokenizers, and training procedures, making behavioral directions difficu

Cross-Sample Relational Fusion: Unifying Domain Generalization and Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2605.08839v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) requires a learning system to learn new classes while retaining previously learned knowledge. However, in real-world sc

CrystalREPA: Transferring Physical Priors from Universal MLIPs to Crystal Generative Models

Model ReleasesDGX agent

arXiv:2605.08960v1 Announce Type: cross Abstract: Crystal generative models mainly learn what stable crystals look like, with little explicit supervision for what makes them stable. We reveal a substa

CT-IDP: Segmentation-Derived Quantitative Phenotypes for Interpretable Abdominal CT Disease Classification

Model ReleasesDGX agent

arXiv:2605.09002v1 Announce Type: cross Abstract: In this retrospective multi-institutional study, a quantitative phenotyping framework, CT-IDP (CT Image-Derived Phenotypes) was developed on the MERLI

CTQWformer: A CTQW-based Transformer for Graph Classification

Model ReleasesDGX agent

arXiv:2605.09486v1 Announce Type: cross Abstract: Graph Neural Networks (GNN) and Transformer-based architectures have achieved remarkable progress in graph learning, yet they still struggle to captur

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging

Model ReleasesDGX agent

arXiv:2605.08455v1 Announce Type: new Abstract: Debugging CUDA programs has long been challenging because failures often arise from subtle interactions among hardware behavior, compiler decisions, mem

CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs

Model ReleasesDGX agent

arXiv:2605.08467v1 Announce Type: new Abstract: Large language models show promise for automated CUDA programming, however even the strongest coding models (e.g., Claude-Opus-4.6) may still fall short

cuRegOT: A GPU-Accelerated Solver for Entropic-Regularized Optimal Transport

Model ReleasesDGX agent

arXiv:2605.08793v1 Announce Type: cross Abstract: Optimal transport (OT) has emerged as a fundamental tool in modern machine learning, yet its computational cost remains a significant bottleneck for l

datasette 1.0a29

Model ReleasesDGX agent

Release: datasette 1.0a29 New TokenRestrictions.abbreviated(datasette) utility method for creating '_r' dictionaries. #2695 Table headers and column options are now visible even if a table contains ze

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices

Model ReleasesDGX agent

arXiv:2605.10933v1 Announce Type: cross Abstract: While Mixture-of-Experts (MoE) scales model capacity without proportionally increasing computation, its massive total parameter footprint creates sign

Decomposing and Steering Functional Metacognition in Large Language Models

Model ReleasesDGX agent

arXiv:2605.08942v1 Announce Type: new Abstract: Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies

Deep learning hit a wall. Neurosymbolic AI rescued it.

Model ReleasesDGX agent

Deep learning hit a wall. Neurosymbolic AI rescued it. 🤩🤯🤩 Claude Code (still not AGI but biggest advance since GPT-4) is the most neurosymbolic thing I have ever seen in my life. 53 symbolic tools, 5

Deep Learning under Fractional-Order Differential Privacy

Model ReleasesDGX agent

arXiv:2605.09890v1 Announce Type: cross Abstract: Differentially private stochastic gradient descent (DP-SGD) is a standard approach to privacy-preserving learning based on per-example clipping, subsa

Deepfake Detection that Generalizes Across Benchmarks

Model ReleasesDGX agent

arXiv:2508.06248v4 Announce Type: replace Abstract: The generalization of deepfake detectors to unseen manipulation techniques remains a challenge for practical deployment. Although many approaches ad

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10564v1 Announce Type: new Abstract: End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual rea

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

Model ReleasesDGX agent

arXiv:2605.09679v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, e

DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos

Model ReleasesDGX agent

arXiv:2605.09586v1 Announce Type: new Abstract: World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and m

Detect, Localize, and Explain: Interactive Hierarchical Log Anomaly Analytics with LLM Augmentation

Model ReleasesDGX agent

arXiv:2605.09222v1 Announce Type: cross Abstract: Logs are ubiquitous in modern systems. Unfortunately, their unstructured nature in flat sequences limits understanding of execution behaviors, hinderi

Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

Model ReleasesDGX agent

arXiv:2604.01151v2 Announce Type: replace Abstract: As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human o

Diagnosing Spectral Ceilings in Equivariant Neural Force Fields

Model ReleasesDGX agent

arXiv:2605.08286v1 Announce Type: cross Abstract: We introduce a spectral-injection diagnostic for measuring which angular frequencies a trained equivariant force-field backbone preserves: inject a co

DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules

Model ReleasesDGX agent

arXiv:2605.08614v1 Announce Type: new Abstract: Monitoring complex industrial assets relies on engineer-authored symbolic rules that trigger based on sensor conditions and prompt technicians to perfor

DiffATS: Diffusion in Aligned Tensor Space

Model ReleasesDGX agent

arXiv:2605.09275v1 Announce Type: new Abstract: Direct diffusion modeling of high-resolution spatiotemporal fields is computationally challenging. Parameter-efficient primitives address this by repres

Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression

Model ReleasesDGX agent

arXiv:2605.08568v1 Announce Type: new Abstract: Large language models (LLMs) have rapidly grown in scale, creating substantial memory and computational costs that hinder efficient deployment. Singular

Diffusion Models are Evolutionary Algorithms

Model ReleasesDGX agent

arXiv:2410.02543v3 Announce Type: replace-cross Abstract: In a convergence of machine learning and biology, we reveal that diffusion models are evolutionary algorithms. By considering evolution as a d

Direct Bethe Free Energy Minimization for Bayesian Neural Ne twork

Model ReleasesDGX agent

arXiv:2605.08446v1 Announce Type: new Abstract: We propose training Bayesian neural networks by directly minimizing the Bethe free energy rather than maximizing a variational lower bound. On tree-stru

Distributional Spectral Diagnostics for Localizing Grokking Transitions

Model ReleasesDGX agent

arXiv:2605.08237v1 Announce Type: new Abstract: In grokking, a model first fits the training data while test accuracy remains low, and only later begins to generalize. We ask whether this transition c

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

Model ReleasesDGX agent

arXiv:2605.08462v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge in Large Language Models (LLMs), particularly in context-grounded settings such as RAG and agentic AI sys

Do Foundation Model Embeddings Improve Cross-Country Crop Yield Generalisation? A Leave-One-Country-Out Evaluation in Sub-Saharan Africa

Model ReleasesDGX agent

arXiv:2605.08113v1 Announce Type: cross Abstract: Accurate predictions of smallholder maize yields across national boundaries are critical for food security planning in sub-Saharan Africa, yet most pu

Do not copy and paste! Rewriting strategies for code retrieval

Model ReleasesDGX agent

arXiv:2605.08299v1 Announce Type: cross Abstract: Embedding-based code retrieval often suffers when encoders overfit to surface syntax. Prior work mitigates this by using LLMs to rephrase queries and

Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation

Model ReleasesDGX agent

arXiv:2605.09315v1 Announce Type: new Abstract: Recent advances in LLM agents enable systems that autonomously refine workflows, accumulate reusable skills, self-train their underlying models, and mai

Do you actually own your document parsing infrastructure? 👀 At @llama_index, we wanted to make that easier, so we built 𝗹𝗶𝘁𝗲𝗽𝗮𝗿𝘀𝗲-…

Model ReleasesDGX agent

Do you actually own your document parsing infrastructure? 👀 At @llama_index, we wanted to make that easier, so we built 𝗹𝗶𝘁𝗲𝗽𝗮𝗿𝘀𝗲-𝘀𝗲𝗿𝘃𝗲𝗿, a lightweight HTTP backend built on top of LiteParse that can

DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

Model ReleasesDGX agent

arXiv:2605.08888v1 Announce Type: new Abstract: Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.08747v1 Announce Type: new Abstract: Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call te

Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces

Model ReleasesDGX agent

arXiv:2605.09497v1 Announce Type: new Abstract: Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Ex

Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval

Model ReleasesDGX agent

arXiv:2504.21015v4 Announce Type: replace-cross Abstract: Training effective dense retrieval models typically relies on hard negative (HN) examples mined from large document corpora using methods such

Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning

Model ReleasesDGX agent

arXiv:2605.09537v1 Announce Type: new Abstract: Despite rapid progress in Vision-Language-Action (VLA) models for robotic control, instruction drift remains a persistent failure mode in long-horizon t

DRNet: All-in-One Image Restoration via Prior-Guided Dynamic Reparameterization

Model ReleasesDGX agent

arXiv:2605.08627v1 Announce Type: new Abstract: All-in-one image restoration aims to handle diverse degradations within a single model. However, existing methods often suffer from three key limitation

DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments

Model ReleasesDGX agent

arXiv:2503.06047v2 Announce Type: replace Abstract: Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent i

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards

Model ReleasesDGX agent

arXiv:2605.08441v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) generates hundreds of thousands of tokens per training step, with rollout generation dominating

Dynamics-Aligned Shared Hypernetworks for Contextual RL under Discontinuous Shifts

Model ReleasesDGX agent

arXiv:2602.06550v2 Announce Type: replace-cross Abstract: Zero-shot generalization in contextual reinforcement learning remains a core challenge, particularly when the context is latent and must be in

Echo-LoRA: Parameter-Efficient Fine-Tuning via Cross-Layer Representation Injection

Model ReleasesDGX agent

arXiv:2605.08177v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) has become a practical route for adapting large language models to downstream tasks, with LoRA-style methods be

EchoAlign: Bridging Generative and Discriminative Learning under Noisy Labels

Model ReleasesDGX agent

arXiv:2405.12969v3 Announce Type: replace Abstract: Noisy labels severely hinder the accuracy and generalization of machine learning models, especially when ambiguous instance features make reliable a

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies

Model ReleasesDGX agent

arXiv:2602.09514v3 Announce Type: replace-cross Abstract: Long-horizon planning is widely recognized as a core capability of autonomous LLM-based agents; however, current evaluation frameworks suffer

EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments

Model ReleasesDGX agent

arXiv:2506.08136v3 Announce Type: replace Abstract: We introduce EconWebArena, a benchmark for evaluating autonomous agents on complex, multimodal economic tasks in realistic web environments. The ben

Edge-specific signal propagation on mature chromophore-region 3D mechanism graphs for fluorescent protein quantum-yield prediction

Model ReleasesDGX agent

arXiv:2605.06644v2 Announce Type: replace Abstract: Fluorescent protein quantum yield (QY) is governed by the mature chromophore and its three-dimensional microenvironment rather than sequence identit

EdgeFlowerTune: Evaluating Federated LLM Fine-Tuning Under Realistic Edge System Constraints

Model ReleasesDGX agent

arXiv:2605.08636v1 Announce Type: new Abstract: Federated fine-tuning offers a promising paradigm for adapting large language models (LLMs) on edge devices by leveraging the rich, diverse, and continu

← Previous
1…261262263264265…377
Next →