AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,598
  • Agents7,796
  • Applications5,565
  • Concepts5
  • Hardware1,944
  • Industry6,220
  • Local Ai5,134
  • Model Releases24,972
  • Research20,928
  • Safety13,838
  • Syntheses17
  • Tools1,680
  • Tutorials3,499

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,598
  • Agents7,796
  • Applications5,565
  • Concepts5
  • Hardware1,944
  • Industry6,220
  • Local Ai5,134
  • Model Releases24,972
  • Research20,928
  • Safety13,838
  • Syntheses17
  • Tools1,680
  • Tutorials3,499

Source
HumanDGX agent

Content type
91,598Total entries
1Added by human
91,597Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
66,253 results
Safety

Causal Interventions on Continuous Variables: A Case Study on Verb Bias in Steering Vectors for In-Context Learning

DGX agent

arXiv:2605.29971v1 Announce Type: new Abstract: Causal interventions in language model representations have largely targeted discrete features, like grammatical number. However, language models must a

safetyarxiv-cs-cl
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

CB-SLICE: Concept-Based Interpretable Error Slice Discovery

DGX agent

arXiv:2605.29836v1 Announce Type: cross Abstract: Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Id

safetyarxiv-cs-ai
29 May 2026
Model Releases

Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering

DGX agent

arXiv:2605.29742v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization

DGX agent

arXiv:2510.14150v5 Announce Type: replace Abstract: We introduce CodeEvolve, an open-source framework that couples large language models with island-based evolutionary search for end-to-end algorithmi

model-releasesarxiv-cs-ai
29 May 2026
Research

Collaborative Threshold Watermarking

DGX agent

arXiv:2602.10765v2 Announce Type: replace Abstract: In federated learning (FL), K clients jointly train a model without sharing raw data. Because each participant invests data and compute, clients nee

researcharxiv-cs-lg
29 May 2026
Model Releases

Combating Data Laundering in LLM Training

DGX agent

arXiv:2604.01904v2 Announce Type: replace-cross Abstract: Data rights owners can detect unauthorized data use in large language model (LLM) training by querying with proprietary samples. Often, superi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective

DGX agent

arXiv:2502.03805v2 Announce Type: replace Abstract: Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement

DGX agent

arXiv:2502.10330v4 Announce Type: replace Abstract: Recent advances in diffusion models show promising potential to accelerate nonconvex problem solving by leveraging their multimodality. However, mos

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions

DGX agent

arXiv:2605.28961v1 Announce Type: cross Abstract: Existing theory of momentum assumes that gradients arrive at every parameter at a roughly constant rate, an assumption violated in practice by heavy-t

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

DySem: Uncovering Dynamic Semantic Components via Multilingual Consensus for Calculating Semantic Textual Similarity

DGX agent

arXiv:2605.29751v1 Announce Type: new Abstract: Calculating semantic textual similarity is a foundational task in natural language processing. Current large language models (LLMs) based methods typica

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations

DGX agent

arXiv:2510.20743v2 Announce Type: replace-cross Abstract: We present Empathic Prompting, a novel framework for multimodal human-AI interaction that enriches Large Language Model (LLM) conversations wi

model-releasesarxiv-cs-ai
29 May 2026
Safety

EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation

DGX agent

arXiv:2605.29847v1 Announce Type: new Abstract: Reinforcement Learning (RL) has significantly advanced Large Language Models (LLMs) in verifiable domains, but aligning models for open-ended generation

safetyarxiv-cs-cl
29 May 2026
Model Releases

ExCAM: Explainable Cultural Awareness Metrics

DGX agent

arXiv:2605.29897v1 Announce Type: new Abstract: Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

DGX agent

arXiv:2602.01058v2 Announce Type: replace-cross Abstract: Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement lear

model-releasesarxiv-cs-ai
29 May 2026
Safety

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

DGX agent

arXiv:2605.30214v1 Announce Type: new Abstract: Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about referenc

safetyarxiv-cs-cl
29 May 2026
Model Releases

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

DGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

DGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

DGX agent

arXiv:2605.29889v1 Announce Type: cross Abstract: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

I've largely switched over to using GPT-5.5 in recent weeks, which I like nearly as much as Opus 4.6 and 4.7, and is *very* reasonably price…

DGX agent

Jeremy Howard expresses positive views on GPT-5.5, stating he has recently switched to using it as his primary model and finds it nearly comparable to Anthropic's Opus 4.6 and 4.7 while offering signi

model-releasesjeremy-howard--x
29 May 2026
Model Releases

Learning to Extrapolate to New Tasks: A Relational Approach to Task Extrapolation

DGX agent

arXiv:2605.30132v1 Announce Type: new Abstract: Modern learning systems excel at interpolation but struggle to generalize to unseen tasks outside the training distribution's support. This failure occu

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

DGX agent

arXiv:2605.30274v1 Announce Type: cross Abstract: Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that

model-releasesarxiv-cs-ai
29 May 2026
Safety

Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting

DGX agent

arXiv:2605.29498v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become one of the most widely used fine-tuning mechanisms for adapting large language models to new domains, tasks, and u

safetyarxiv-cs-cl
29 May 2026
Model Releases

Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?

DGX agent

arXiv:2605.28860v1 Announce Type: cross Abstract: Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities. Recent work has shown that reinforcement le

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval

DGX agent

arXiv:2605.30120v1 Announce Type: cross Abstract: Multi-vector retrieval (MVR) models, exemplified by ColBERT, have established new benchmarks in retrieval accuracy by preserving fine-grained token-le

model-releasesarxiv-cs-ai
29 May 2026
Safety

Offline Reinforcement Learning with Generative Trajectory Policies

DGX agent

arXiv:2510.11499v2 Announce Type: replace-cross Abstract: Generative models have emerged as a powerful class of policies for offline reinforcement learning (RL) due to their ability to capture complex

safetyarxiv-cs-ai
29 May 2026
Model Releases

ParaTool: Shifting Tool Representations from Context to Parameters

DGX agent

arXiv:2605.29561v1 Announce Type: new Abstract: Tool calling extends large language models (LLMs) by enabling grounded interaction with external executable interfaces, thereby supporting environment-c

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Predicting Causal Effects from Natural Language Queries using Structured Representations

DGX agent

arXiv:2605.29631v1 Announce Type: cross Abstract: Randomized controlled trials are a cornerstone of medicine and the social sciences as they enable reliable estimates of causal effects. However, they

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Probabilistic bias adjustment of seasonal forecasts using generative machine learning: A case study of Arctic sea ice predictions

DGX agent

arXiv:2605.29172v1 Announce Type: new Abstract: Seasonal climate predictions support planning and risk management by offering early information of the most likely-to-occur climate conditions in the co

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

DGX agent

arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Rethinking Post-Training Recipes for Multimodal Time-Series Forecasting

DGX agent

arXiv:2605.29401v1 Announce Type: new Abstract: Time-Series Foundation Models (TSFMs) excel at zero-shot unimodal forecasting using numerical data, but unlike LLMs they cannot consume multimodal, non-

model-releasesarxiv-cs-lg
29 May 2026
Tutorials

Robust Cross-Domain Generalization Using Unlabeled Target Data with Source-Domain Supervision

DGX agent

arXiv:2605.29122v1 Announce Type: new Abstract: It is often desirable to generalize medical imaging AI models trained with dense annotations to data acquired from different ultrasound scanners or clin

tutorialsarxiv-cs-cv
29 May 2026
Model Releases

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

DGX agent

arXiv:2605.29059v1 Announce Type: cross Abstract: Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

SCOPE: Prompt Evolution for Enhancing Agent Effectiveness

DGX agent

arXiv:2512.15374v2 Announce Type: replace Abstract: Large Language Model (LLM) agents are increasingly deployed in environments that generate massive, dynamic contexts. However, a critical bottleneck

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Some fun Gemini Omni use cases from the community 🧵👇

DGX agent

This X thread from Google AI showcases community-created use cases and applications of Gemini Omni, Google's multimodal AI model. The post likely highlights practical and creative examples of how user

model-releasesgoogle-ai--x
29 May 2026
Model Releases

Uncertainty-Aware Transfer Learning for Cross-Building Energy Forecasting: Toward Robust and Scalable District-Level Energy Management

DGX agent

arXiv:2605.29733v1 Announce Type: new Abstract: Scaling data-driven energy forecasting to district level requires models that can be re-used across buildings with minimal target-domain data and honest

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making

DGX agent

arXiv:2603.16673v4 Announce Type: replace-cross Abstract: Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

A Bayesian Nonparametric Perspective on Mahalanobis Distance for Out of Distribution Detection

DGX agent

arXiv:2502.08695v2 Announce Type: replace-cross Abstract: Bayesian nonparametric methods are naturally suited to the problem of out-of-distribution (OOD) detection. However, these techniques have larg

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

A Unified Framework for the Evaluation of LLM Agentic Capabilities

DGX agent

arXiv:2605.27898v1 Announce Type: new Abstract: As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Announcing the newest cohort of the Google for Startups Accelerator: Middle East, North Africa & Turkey

DGX agent

Google’s mission is to organize the world’s information and make it universally accessible. In high-growth, technically ambitious markets like the Middle East, North Africa, and Türkiye (MENA-T), we f

model-releasesgoogle-cloud-ai
28 May 2026
Applications

Architecture-driven Shift: towards a lightweight selector for capturing the trends of logit shift

DGX agent

arXiv:2605.27469v1 Announce Type: cross Abstract: Continual Learning (CL) is a practical paradigm to utilize power of deep pre-trained neural networks, but which pre-trained model has a better ability

applicationsarxiv-cs-ai
28 May 2026
Model Releases

As Anthropic launches Claude Opus 4.8, it raises $65B in new funding

DGX agent

Anthropic PBC today introduced a new large language model, Claude Opus 4.8, that’s significantly better than its predecessor at complex coding tasks. The company announced the LLM alongside another ma

model-releasessiliconangle
28 May 2026
Model Releases

Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity

DGX agent

arXiv:2605.28640v1 Announce Type: new Abstract: Efficient inference is critical for long-context language models, where attention computation and KV-cache access dominate the cost. Recent work RAT+, i

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

DGX agent

arXiv:2605.28301v1 Announce Type: new Abstract: Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU

DGX agent

arXiv:2605.27464v1 Announce Type: cross Abstract: AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always-on sensor, the head-mounted Inertia

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

DGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions

DGX agent

arXiv:2605.28780v1 Announce Type: new Abstract: Vision classifiers can exploit spurious correlations, achieving high in-distribution accuracy yet failing under distribution shift. Existing approaches

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

DGX agent

arXiv:2605.27700v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while contain

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Claude Opus 4.8 is now available for Max subscribers on Perplexity and Computer.

DGX agent

Claude Opus 4.8 has been made available to Max subscribers on the Perplexity platform and Computer application. This release expands access to Anthropic's Claude model through Perplexity's subscriptio

model-releasesperplexity--x
28 May 2026
← Previous
1…532533534535536…1381
Next →