AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
91,060Total entries
1Added by human
91,059Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,793 results
Safety

ClaHF: A Human Feedback-inspired Reinforcement Learning Framework for Improving Classification Tasks

DGX agent

arXiv:2605.17458v1 Announce Type: new Abstract: Text classification models are typically trained via supervised fine-tuning (SFT). However, SFT essentially performs behavior cloning from instance-wise

safetyarxiv-cs-lg
19 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Closing the Gap at CRAC 2026: Two-Stage Adaptation for LLM-Based Multilingual Coreference Resolution

DGX agent

arXiv:2605.16984v1 Announce Type: new Abstract: We present our submission to the LLM track of the 2026 Computational Models of Reference, Anaphora and Coreference (CRAC 2026) shared task. With an aver

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

ContractBench: Can LLM Agents Preserve Observation Contracts?

DGX agent

arXiv:2605.17281v1 Announce Type: cross Abstract: Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation co

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning

DGX agent

arXiv:2602.02979v2 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong potential in complex reasoning, yet their progress remains fundamentally constrained by reliance

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs

DGX agent

arXiv:2605.18498v1 Announce Type: cross Abstract: Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balanci

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Disappointing pricing trend with Gemini 3.5 Flash. 22.5x pricier than 2.0 Flash which came out 15 months ago (9.00 vs 0.40). Are Flash mod…

DGX agent

Disappointing pricing trend with Gemini 3.5 Flash. 22.5x pricier than 2.0 Flash which came out 15 months ago (9.00 vs 0.40). Are Flash models supposed to get this much more expensive, or is Pro just b

model-releasesjeremy-howard--x
19 May 2026
Model Releases

Distributed Perceptron under Bounded Staleness, Partial Participation, and Noisy Communication

DGX agent

arXiv:2601.10705v3 Announce Type: replace Abstract: We study a semi-asynchronous client-server perceptron trained via iterative parameter mixing (IPM-style averaging): clients run local perceptron upd

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

DriveSafer: End-to-End Autonomous Driving with Safety Guidance

DGX agent

arXiv:2605.16737v1 Announce Type: cross Abstract: End-to-End (E2E) autonomous driving models have shown growing capability in recent years, with performance improving on increasingly challenging bench

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

DynMuon: A Dynamic Spectral Shaping View of Muon

DGX agent

arXiv:2605.17109v1 Announce Type: cross Abstract: In recent years, Muon has emerged as the dominant method for training large language models, and transformers more broadly. The essential difference,

model-releasesarxiv-cs-ai
19 May 2026
Research

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks

DGX agent

arXiv:2603.11689v2 Announce Type: replace Abstract: Frontier Multimodal Large Language Models (MLLMs) exhibit remarkable capabilities in Visual-Language Comprehension (VLC) tasks. However, they are of

researcharxiv-cs-ai
19 May 2026
Local Ai

Federated Learning by Utility-Constrained Stochastic Aggregation for Improving Rational Participation

DGX agent

arXiv:2605.18020v1 Announce Type: new Abstract: Federated Learning (FL) algorithms implicitly assume that clients passively comply with server-side orchestration by sharing local model updates upon se

local-aiarxiv-cs-lg
19 May 2026
Local Ai

FedSDR: Federated Self-Distillation with Rectification

DGX agent

arXiv:2605.18028v1 Announce Type: cross Abstract: Federated fine-tuning of Large Language Models faces severe statistical heterogeneity. However, existing model-level defenses often overlook the root

local-aiarxiv-cs-ai
19 May 2026
Model Releases

Gemini 3.5 Flash is now available in Windsurf!

DGX agent

Cognition AI announced the availability of Gemini 3.5 Flash within Windsurf, their AI development environment. This integration brings Google's faster Gemini 3.5 Flash model to Windsurf users, likely

model-releasescognition-ai--x
19 May 2026
Research

Haptic Rendering of Fractional-Order Viscoelasticity: Passivity and Rendering Fidelity

DGX agent

arXiv:2605.16389v1 Announce Type: cross Abstract: Haptic rendering of viscoelastic materials that exhibit creep and stress relaxation is crucial for many applications, such as medical training with re

researcharxiv-cs-ai
19 May 2026
Model Releases

High-dimensional ridge regression with random features for non-identically distributed data with a variance profile

DGX agent

arXiv:2504.03035v2 Announce Type: replace-cross Abstract: Random feature ridge regression is often analyzed in the high-dimensional regime under the homogeneous sampling model x_i=Sigma^{1/2}x_i', whe

model-releasesarxiv-cs-lg
19 May 2026
Local Ai

In-context learning enables continental-scale subsurface temperature prediction from sparse local observations

DGX agent

arXiv:2605.16665v1 Announce Type: new Abstract: Continental-scale knowledge of subsurface temperature is limited by the cost and sparsity of borehole measurements, but such information is essential fo

local-aiarxiv-cs-lg
19 May 2026
Model Releases

Joint Parameter and State-Space Bayesian Optimization: Using Process Expertise to Accelerate Manufacturing Optimization

DGX agent

arXiv:2602.17679v2 Announce Type: replace Abstract: Bayesian optimization (BO) is a powerful method for optimizing black-box manufacturing processes, but its performance is often limited when dealing

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

Latency-Aware Deep Learning Benchmark for Real-Time Cyber-Physical Attack and Fault Classification in Inverter-Dominated Power Grids

DGX agent

arXiv:2605.17256v1 Announce Type: cross Abstract: This work introduces a latency-aware benchmarking framework for evaluating deep learning models in power system anomaly detection using high-fidelity,

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

LEAF: A Living Benchmark for Event-Augmented Forecasting

DGX agent

arXiv:2605.16358v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly applied to forecasting. To evaluate this capability while mitigating pre-training data contamination, se

model-releasesarxiv-cs-ai
19 May 2026
Research

Learning more physically realistic dynamics in machine-learning based weather forecasting with latent-space constraints

DGX agent

arXiv:2510.04006v2 Announce Type: replace Abstract: Data-driven machine learning (ML) models are reshaping weather forecasting and have shown the potential to accelerate and surpass traditional physic

researcharxiv-cs-lg
19 May 2026
Research

Learning Quantifiable Visual Explanations Without Ground-Truth

DGX agent

arXiv:2605.18681v1 Announce Type: new Abstract: Explainable AI (XAI) techniques are increasingly important for the validation and responsible use of modern deep learning models, but are difficult to e

researcharxiv-cs-ai
19 May 2026
Model Releases

LoopQ: Quantization for Recursive Transformers

DGX agent

arXiv:2605.16343v1 Announce Type: cross Abstract: Looped language models (LoopLMs) improve parameter efficiency by recursively reusing Transformer blocks, enabling deeper computation under a fixed mod

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation

DGX agent

arXiv:2605.17292v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems have shown promise for solving complex tasks through agent collaboration. However, existing frameworks as

model-releasesarxiv-cs-ai
19 May 2026
Safety

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

DGX agent

arXiv:2605.18549v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not alwa

safetyarxiv-cs-cl
19 May 2026
Model Releases

My notes on Gemini 3.5 Flash - 3x the price of Gemini 3 Flash but Google are planning to use it for many of their own products https://simon…

DGX agent

Google's Gemini 3.5 Flash model costs approximately 3x more than Gemini 3 Flash, despite being a newer version. Google plans to integrate Gemini 3.5 Flash into many of their own products, suggesting t

model-releasessimon-willison--x
19 May 2026
Model Releases

NeuSymMS: A Hybrid Neuro-Symbolic Memory System for Persistent, Self-Curating LLM Agents

DGX agent

arXiv:2605.17596v1 Announce Type: new Abstract: We present NeuSymMS, an adaptive memory system that enables large language model (LLM) agents to learn, remember, and reason about users across sessions

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks

DGX agent

arXiv:2605.18583v1 Announce Type: cross Abstract: Coding agents now run autonomously with shell, file, and network privileges. When a user issues a benign request, the agent sometimes does more than a

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

PaliBench: A Multi-Reference Blueprint for Classical Language Translation Benchmarks

DGX agent

arXiv:2605.16881v1 Announce Type: new Abstract: Digital humanities projects increasingly rely on machine translation and large language models to widen access to classical, religious, and otherwise un

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments

DGX agent

arXiv:2603.23231v2 Announce Type: replace Abstract: Empowering large language models with long-term memory is crucial for building agents that adapt to users' evolving needs. Existing evaluations of t

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media

DGX agent

arXiv:2605.17187v1 Announce Type: cross Abstract: Social media are shifting towards pluralism -- community-governed platforms where groups define their own norms. What violates rules in one community

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction

DGX agent

arXiv:2605.18053v1 Announce Type: cross Abstract: We study KV cache eviction under a shared globally capped decode-time harness. Seven policies (LRU, H2O, SnapKV, StreamingLLM, Ada-KV, QUEST, Random)

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification

DGX agent

arXiv:2510.12534v4 Announce Type: replace Abstract: The rapid growth of user-generated text across digital platforms has intensified the need for interpretable models capable of fine-grained text clas

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Reference anything: Gemini Omni extends Gemini's native multimodality, allowing you to blend combinations of text, audio, image, and video i…

DGX agent

Gemini Omni is an extension of Google's Gemini model that enhances its multimodal capabilities by enabling seamless integration of text, audio, image, and video inputs and outputs. This advancement al

model-releasesgoogle-ai--x
19 May 2026
Model Releases

SAM 2++: Tracking Anything at Any Granularity

DGX agent

arXiv:2510.18822v4 Announce Type: replace Abstract: Due to the varying granularity of target states across different tasks, most existing trackers are tailored to a single task, which specificity limi

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection

DGX agent

arXiv:2605.17311v1 Announce Type: new Abstract: The remarkable visual fidelity of recent commercial video generative models, such as Sora and Veo, renders robust AI-generated video detection increasin

model-releasesarxiv-cs-cv
19 May 2026
Hardware

Stable Audio 3

DGX agent

arXiv:2605.17991v1 Announce Type: cross Abstract: Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models c

hardwarearxiv-cs-ai
19 May 2026
Tutorials

StAD: Stein Amortized Divergence for Fast Likelihoods with Diffusion and Flow

DGX agent

arXiv:2605.16486v1 Announce Type: cross Abstract: Diffusion and flow-based models are ubiquitously used for generative modelling and density estimation. They admit a deterministic probability flow ord

tutorialsarxiv-cs-lg
19 May 2026
Model Releases

State-of-the-Art Claims Require State-of-the-Art Evidence

DGX agent

arXiv:2605.17273v1 Announce Type: cross Abstract: State-of-the-Art (SOTA) claims pervade Artificial Intelligence (AI) and Machine Learning (ML) research. These claims rest on benchmark evaluations, wh

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Supervising the search process produces reliable and generalizable information-seeking agents

DGX agent

arXiv:2502.13957v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming web search by shifting from document ranking to synthesizing answers, and are increasingly deplo

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning

DGX agent

arXiv:2605.18109v1 Announce Type: new Abstract: In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints

DGX agent

arXiv:2602.21265v2 Announce Type: replace Abstract: We introduce ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. ToolMAT

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey

DGX agent

arXiv:2409.10102v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has quickly grown into a pivotal paradigm in the development of Large Language Models (LLMs). Although ex

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

TusoAI: Agentic Optimization for Scientific Methods

DGX agent

arXiv:2509.23986v2 Announce Type: replace Abstract: Scientific discovery is often slowed by the manual development of computational tools needed to analyze complex experimental data. Building such too

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

DGX agent

arXiv:2605.18740v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evide

local-aiarxiv-cs-ai
19 May 2026
Model Releases

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

DGX agent

arXiv:2605.18172v1 Announce Type: new Abstract: Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering

DGX agent

arXiv:2605.18313v1 Announce Type: cross Abstract: Small vision-language models (2-8B) are well-suited for clin- ical deployment due to privacy constraints, limited connectivity, and low-latency requir

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

DGX agent

arXiv:2601.17887v2 Announce Type: replace Abstract: Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized ag

model-releasesarxiv-cs-ai
19 May 2026
Safety

Why Do Safety Guardrails Degrade Across Languages?

DGX agent

arXiv:2605.17173v1 Announce Type: cross Abstract: Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds

safetyarxiv-cs-ai
19 May 2026
← Previous
1…465466467468469…1371
Next →