AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

AgenticDataBench: A Comprehensive Benchmark for Data Agents

DGX agent

arXiv:2607.01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern so

model-releasesarxiv-cs-ai
3 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG

DGX agent

arXiv:2602.19127v2 Announce Type: replace Abstract: With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop rea

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

DGX agent

arXiv:2607.02255v1 Announce Type: new Abstract: Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see. The simplest contract appends past observations, to

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

DGX agent

arXiv:2607.01934v1 Announce Type: cross Abstract: This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk asses

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms

DGX agent

arXiv:2603.29466v2 Announce Type: replace-cross Abstract: Existing methods for quantifying predictive uncertainty in neural networks are either computationally intractable for large language models or

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

DGX agent

arXiv:2607.02269v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are l

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Ask the Right Comparison:Bias-Aware Bayesian Active Top-k Ranking with LLM Judges

DGX agent

arXiv:2607.02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

DGX agent

arXiv:2607.01973v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answerin

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

DGX agent

arXiv:2607.02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking diffi

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

BALF: Budgeted Activation-Aware Low-Rank Factorization for Fine-Tuning-Free Model Compression

DGX agent

arXiv:2509.25136v3 Announce Type: replace Abstract: Activation-aware low-rank factorization techniques yield strong compression results but are generally confined to linear layers, while existing whit

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation

DGX agent

arXiv:2607.02182v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit remarkable reasoning capabilities, but their task-specific fine-tuning is notoriously plagued by overconfidence,

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Benchmarking Federated Learning and Knowledge Distillation for Point Cloud Classification

DGX agent

arXiv:2607.01272v1 Announce Type: cross Abstract: Deploying 3D point cloud analysis in privacy-sensitive, resource-constrained settings faces two barriers: data cannot be centralized, and models must

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Beyond Pixel Diffs: Benchmarking Image Change Captioning for Web UI Visual Regression Testing

DGX agent

arXiv:2607.01728v1 Announce Type: cross Abstract: Visual regression testing (VRT) is a standard quality assurance step in modern software release pipelines. On every change, it re-renders user interfa

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework

DGX agent

arXiv:2607.01581v1 Announce Type: new Abstract: The capacity of Large Language Models (LLMs) to reason about pedagogical intent within instructional communication remains underexplored, particularly i

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Black-Box Inference of LLM Architectural Properties with Restrictive API Access

DGX agent

arXiv:2607.01313v1 Announce Type: cross Abstract: In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures. However, prior work has shown that given l

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Born Discrete, Made Smooth: Variational Formulation of Shallow Neural Networks

DGX agent

arXiv:2607.02003v1 Announce Type: cross Abstract: Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-c

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Boundary-Aware Quantization: Finite-Scale Decision Geometry of Neural Classifiers

DGX agent

arXiv:2607.01478v1 Announce Type: cross Abstract: We measured quantization-induced decision-boundary changes using local logit-margin radii, first-order boundary displacement, normal variation, slice-

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systems

DGX agent

arXiv:2607.01600v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed as communicating agents, does inter-agent communication cause outputs to converge? We introduce BOUNDARY_

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment

DGX agent

arXiv:2607.01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

BRIDGE: Predicting Human Task Completion Time From Model Performance

DGX agent

arXiv:2602.07267v2 Announce Type: replace Abstract: Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Ex

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Bringing Agentic Search to Earth Observation Data Discovery

DGX agent

arXiv:2607.02387v1 Announce Type: cross Abstract: NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

BuilderBench: The Building Blocks of Intelligent Agents

DGX agent

arXiv:2510.06288v4 Announce Type: replace Abstract: Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set b

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation

DGX agent

arXiv:2607.02222v1 Announce Type: cross Abstract: Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decompositio

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows

DGX agent

arXiv:2607.01709v1 Announce Type: new Abstract: Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently. As these workflows become repeated

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Composite Reward Design in PPO-Driven Adaptive Filtering

DGX agent

arXiv:2506.06323v2 Announce Type: replace-cross Abstract: Model-free and reinforcement learning-based adaptive filtering methods are gaining traction for denoising in dynamic, non-stationary environme

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair

DGX agent

arXiv:2607.01916v1 Announce Type: new Abstract: Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Controllable Sim Agents with Behavior Latents

DGX agent

arXiv:2607.02496v1 Announce Type: cross Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enabl

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection

DGX agent

arXiv:2607.01303v1 Announce Type: cross Abstract: Presentation Attack Detection (PAD) serves as a crucial safeguard for face recognition systems against presentation attacks such as printed photos, re

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

DecompRL: Solving Harder Problems by Learning Modular Code Generation

DGX agent

arXiv:2607.02390v1 Announce Type: new Abstract: How can Large Language Models (LLMs) solve problems they currently cannot? Repeated sampling scales test-time compute but GPU cost grows linearly with a

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Decoupling Code Complexity from Newcomer Participation: A Causal Study of AI Coding Agent Adoption in OSS

DGX agent

arXiv:2607.01810v1 Announce Type: cross Abstract: Open-source projects depend on a steady inflow of newcomers. A growing concern is that AI coding agents (tools such as Cursor and Claude Code that wri

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Dendritic In-Context Learning in a Single-Layer Spiking Neural Network

DGX agent

arXiv:2607.02283v1 Announce Type: cross Abstract: In-context learning (ICL) operates via implicit gradient descent embedded in the forward pass of modern AI architectures -- Transformers, Mamba, state

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Denser neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training

DGX agent

arXiv:2607.01763v1 Announce Type: cross Abstract: Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Discrete Diffusion Language Models for Interactive Radiology Report Drafting

DGX agent

arXiv:2607.01436v1 Announce Type: new Abstract: Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have become competi

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Distributed Attacks in Persistent-State AI Control

DGX agent

arXiv:2607.02514v1 Announce Type: new Abstract: As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence crea

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Distributionally Robust Listwise Preference Optimization

DGX agent

arXiv:2607.01715v1 Announce Type: new Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, o

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry

DGX agent

arXiv:2607.01661v1 Announce Type: new Abstract: Multi-agent systems are increasingly used for forecasting future events, as deliberation among multiple LLMs is believed to improve reasoning and calibr

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

DL-VINS-Factory: A Modular Framework for Learned Visual Front-Ends in Visual-Inertial SLAM

DGX agent

arXiv:2607.01757v1 Announce Type: cross Abstract: Deep-learning features excel in visual matching, yet their practical value in tightly coupled visual-inertial SLAM (VI-SLAM) remains insufficiently ch

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

Do Newer Lightweight CNNs Perform Better Under Resource Constraints? A Controlled Multigenerational Study of Architecture, Initialization, Training Budget, and Efficiency

DGX agent

arXiv:2607.01984v1 Announce Type: cross Abstract: Newer lightweight convolutional neural networks are often presented as improving predictive performance and deployment efficiency, but such claims req

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

eCream-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

DGX agent

arXiv:2606.12569v2 Announce Type: replace-cross Abstract: We present eCream-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals. Th

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models

DGX agent

arXiv:2607.02007v1 Announce Type: new Abstract: Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single dis

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

DGX agent

arXiv:2607.02501v1 Announce Type: new Abstract: Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation

DGX agent

arXiv:2607.01584v1 Announce Type: new Abstract: Large language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literature and free-f

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning

DGX agent

arXiv:2607.01789v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models scale efficiently but remain costly to adapt due to redundant experts and uniform parameter allocation. Existing param

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Evidence-State Rewards for Long-Context Reasoning

DGX agent

arXiv:2607.02073v1 Announce Type: new Abstract: Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context RL methods us

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

DGX agent

arXiv:2607.02440v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

eXact-Prior Variational Autoencoder (X-VAE): Learning Data-Adaptive Gaussian Mixture Priors for Latent Distributions

DGX agent

arXiv:2607.01275v1 Announce Type: cross Abstract: Variational Autoencoders (VAEs) commonly assume a standard isotropic Gaussian prior over the latent space, an assumption that often fails to capture t

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability

DGX agent

arXiv:2607.01799v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) decompose internal activations of neural networks into sparse linear combinations of learned features by fitting an overcom

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

DGX agent

arXiv:2607.02396v1 Announce Type: new Abstract: Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behav

model-releasesarxiv-cs-ai
3 Jul 2026
← Previous
1…9293949596…361
Next →