AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

DGX agent

arXiv:2601.22025v2 Announce Type: replace-cross Abstract: Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantically

model-releasesarxiv-cs-ai
11 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

When Roleplaying, Do Models Believe What They Say?

DGX agent

arXiv:2606.11502v1 Announce Type: cross Abstract: Language models can state that 'the Earth orbits the Sun' and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adopt

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs

DGX agent

arXiv:2606.12385v1 Announce Type: new Abstract: Modern LLM training pipelines increasingly rely on other models to generate data, filter corpora, judge outputs, and guide development decisions. These

model-releasesarxiv-cs-cl
11 Jun 2026
Model Releases

World Model Self-Distillation: Training World Models to Solve General Tasks

DGX agent

arXiv:2606.12072v1 Announce Type: new Abstract: Pretrained video generators are promising visual world models that exhibit emergent task-solving abilities; however, their reliance on detailed textual

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

World Pilot: Steering Vision-Language-Action Models with World-Action Priors

DGX agent

arXiv:2606.12403v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models inherit semantic grounding from large-scale pretraining and perform competently across in-distribution manipulation

model-releasesarxiv-cs-ro
11 Jun 2026
Model Releases

WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning

DGX agent

arXiv:2606.11816v1 Announce Type: cross Abstract: Forecasting real-world events requires language-model agents to reason under uncertainty from incomplete, time-bounded information. Yet evaluating whe

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

5% > 100%: Flatness Preference is All You Need for Multimodal Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2606.10488v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods provide a streamlined and efficient tool for adapting large models to domain-specific multimodal downstre

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

A complementary study on PlanGPT: Evaluation with defined Performance Metrics and comparison with a planner

DGX agent

arXiv:2606.10489v1 Announce Type: new Abstract: Automated Planning is a subfield of Artificial Intelligence (AI) where the main objective is generating a sequence of actions, known as a plan, that hel

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS

DGX agent

arXiv:2606.10928v1 Announce Type: cross Abstract: Large language models can reduce the manual effort required to set up finite element simulations, but they introduce reliability risks when generated

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

A History-Aware Visually Grounded Critic for Computer Use Agents

DGX agent

arXiv:2606.11078v1 Announce Type: new Abstract: Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-executio

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

A Large Scale Open-Source Image and Video Dataset for Robust Wildfire Detection and Classification

DGX agent

arXiv:2606.10174v1 Announce Type: new Abstract: Wildfire detection and monitoring are critical for mitigating fire spread and reducing environmental and infrastructural damage. In this work, we introd

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI

DGX agent

arXiv:2505.01458v2 Announce Type: replace-cross Abstract: Navigation and manipulation are core capabilities in Embodied AI, but training agents to perform them directly in the real world is costly, ti

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

DGX agent

arXiv:2606.11150v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experime

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design

DGX agent

arXiv:2606.10493v1 Announce Type: cross Abstract: Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even under low-conc

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

DGX agent

arXiv:2502.11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

Advancing the State-of-the-Art in Empirical Privacy Auditing

DGX agent

arXiv:2606.10481v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning of large language models (LLMs) can exhibit problematic memorization of individual training examples. Empirical privac

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis

DGX agent

arXiv:2606.10381v1 Announce Type: cross Abstract: Muon collider research spans accelerator physics, detector instrumentation, and high-energy phenomenology, with relevant evidence scattered across a r

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness

DGX agent

arXiv:2606.10577v1 Announce Type: new Abstract: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Howe

model-releasesarxiv-cs-ro
10 Jun 2026
Model Releases

AgniNav: Configuration-Driven Cross-Embodiment Local Planning for Robot Navigation

DGX agent

arXiv:2606.10903v1 Announce Type: new Abstract: Monocular local navigation is attractive for lightweight robots, but existing vision-based policies often couple perception to a specific body, camera h

model-releasesarxiv-cs-ro
10 Jun 2026
Model Releases

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

DGX agent

arXiv:2606.09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

An adaptive framework for the axisymmetric pulsar magnetosphere using physics-informed Kolmogorov-Arnold networks

DGX agent

arXiv:2606.10686v1 Announce Type: cross Abstract: The pulsar magnetosphere has only recently been addressed using Physics-Informed Neural Networks (PINNs), by deploying a domain-decomposition approach

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs

DGX agent

arXiv:2603.14463v2 Announce Type: replace Abstract: Adapting Large Language Models (LLMs) to high-stakes vertical domains like insurance presents a significant challenge: scenarios demand strict adher

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Anomaly Detection and Root Cause Analysis for Microservice Systems

DGX agent

arXiv:2606.09942v1 Announce Type: cross Abstract: Microservice systems are widely used to build cloud applications, yet their complexity makes failures inevitable, degrading user experience and causin

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents

DGX agent

arXiv:2602.04935v3 Announce Type: replace-cross Abstract: Adapting LLM agents to domain-specific tool calling remains notably brittle under evolving interfaces. Prompt and schema engineering is easy t

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Assessing Automated Prompt Injection Attacks in Agentic Environments

DGX agent

arXiv:2606.10525v1 Announce Type: cross Abstract: Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effec

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

DGX agent

arXiv:2505.23851v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization wi

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It

DGX agent

arXiv:2606.11052v1 Announce Type: new Abstract: Chain-of-thought (CoT) supervised fine-tuning (SFT) is widely adopted to improve reasoning ability, yet we find that it systematically degrades long-con

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings

DGX agent

arXiv:2606.10716v1 Announce Type: cross Abstract: Pre-trained language models (PLMs) have achieved strong performance in keyphrase extraction (KPE), largely due to their ability to generate rich conte

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

au-Rec: A Verifiable Benchmark for Agentic Recommender Systems

DGX agent

arXiv:2606.10156v1 Announce Type: cross Abstract: As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benc

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

DGX agent

arXiv:2407.20242v5 Announce Type: replace-cross Abstract: Embodied AI represents systems where AI is integrated into physical entities. Large Language Model (LLM), which exhibits powerful language und

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations

DGX agent

arXiv:2606.10281v1 Announce Type: cross Abstract: This paper presents AuditBench, a new benchmark dataset for evaluating the capabilities of LLMs at investigating security-related system audit logs. W

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Benchmarking Knowledge Editing using Logical Rules

DGX agent

arXiv:2606.10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLM

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Benchmarking stereo reconstruction for 3D printable Martian terrain models

DGX agent

arXiv:2606.10364v1 Announce Type: new Abstract: Reconstructing printable 3D models from Mars rover imagery is challenging because Martian terrain is low-texture, irregular, and partially observed. We

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts

DGX agent

arXiv:2606.10061v1 Announce Type: new Abstract: Large language models (LLMs) increasingly participate in emotionally sensitive social conversations, where responses may shift from balanced support tow

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

DGX agent

arXiv:2606.10803v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the 'brain' of embodied AI, instructing robots to i

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Beyond Memorization: Distinguishing Between Pattern-Based and Epistemic Reasoning in LLMs Using Epistemic Puzzles

DGX agent

arXiv:2603.21350v2 Announce Type: replace Abstract: Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior w

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

DGX agent

arXiv:2606.10905v1 Announce Type: new Abstract: Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces

DGX agent

arXiv:2606.10064v1 Announce Type: cross Abstract: Small-model agentic post-training is bottlenecked less by the algorithm than by the trajectory substrate it consumes. Leading recipes (RLVR, group-rel

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Boosting Graph Robustness Against Backdoor Attacks: An Over-Similarity Perspective

DGX agent

arXiv:2502.01272v3 Announce Type: replace Abstract: Graph Neural Networks (GNNs) have achieved notable success in tasks such as social and transportation networks. However, recent studies have highlig

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

CableRobotGraphSim: A Graph Neural Network for Modeling Partially Observable Cable-Driven Robot Dynamics

DGX agent

arXiv:2602.21331v2 Announce Type: replace Abstract: General-purpose simulators have accelerated the development of robots. Traditional simulators based on first-principles, however, typically require

model-releasesarxiv-cs-ro
10 Jun 2026
Model Releases

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency

DGX agent

arXiv:2606.10620v1 Announce Type: cross Abstract: Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly und

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Can Multi-Agent LLMs Identify Their Peers? Stylometric Fingerprinting in Role-Constrained Political Analysis

DGX agent

arXiv:2606.09854v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) pipelines for political statement analysis are vulnerable to peer-preservation bias: models tend to protect pee

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement

DGX agent

arXiv:2606.10640v1 Announce Type: new Abstract: In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structure

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

DGX agent

arXiv:2606.11063v1 Announce Type: new Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This part

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

CITRAS-FM: Tiny Time Series Foundation Model for Covariate-Informed Zero-Shot Forecasting

DGX agent

arXiv:2606.10798v1 Announce Type: new Abstract: Pretrained time series foundation models (TSFMs) have enabled zero-shot forecasting on unseen target series. However, existing TSFMs often incur high co

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

CleanPatrick: A Benchmark for Image Data Cleaning

DGX agent

arXiv:2505.11034v2 Announce Type: replace-cross Abstract: Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, lim

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference

DGX agent

arXiv:2606.10935v1 Announce Type: cross Abstract: Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass. Multi-token prediction (MTP)

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ClusBench: The Clustering Benchmark Data Resource You've All Been Waiting For (?)

DGX agent

arXiv:2606.10673v1 Announce Type: cross Abstract: Although some very common test beds exist for assessing the performance of clustering methods, large scale benchmarking is typically limited to relati

model-releasesarxiv-cs-lg
10 Jun 2026
← Previous
1…135136137138139…361
Next →