AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlog
91,020Total entries
1Added by human
91,019Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,767 results
Research

PEEK: Picking Essential frames via Efficient Knowledge distillation

DGX agent

arXiv:2605.31029v1 Announce Type: new Abstract: Video-language models can process only a limited number of frames, making frame selection a key bottleneck for efficient video captioning. Most captioni

researcharxiv-cs-cv
1 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Plain Transformers are Surprisingly Powerful Link Predictors

DGX agent

arXiv:2602.01553v2 Announce Type: replace-cross Abstract: Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies. While

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

Preference-Aware Rubric Learning for Personalized Evaluation

DGX agent

arXiv:2605.31545v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model beha

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

Probabilistic Precipitation Nowcasting with Rectified Flow Transformers

DGX agent

arXiv:2605.31204v1 Announce Type: new Abstract: Accurate weather forecasts are essential across various domains and are safety-critical in extreme weather conditions. Compared to simulation-based fore

model-releasesarxiv-cs-cv
1 Jun 2026
Safety

Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

DGX agent

arXiv:2504.11972v3 Announce Type: replace Abstract: Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Rece

safetyarxiv-cs-cl
1 Jun 2026
Safety

Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

DGX agent

arXiv:2412.03876v2 Announce Type: replace Abstract: Text-to-Image (T2I) diffusion models are widely recognized for their ability to generate high-quality and diverse images based on text prompts. Howe

safetyarxiv-cs-cv
1 Jun 2026
Research

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

DGX agent

arXiv:2605.31433v1 Announce Type: new Abstract: Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dep

researcharxiv-cs-cl
1 Jun 2026
Applications

Spatio-temporal stochastic graph-based learning for infectious disease forecasting

DGX agent

arXiv:2605.30662v1 Announce Type: new Abstract: Spatio-temporal graph-based models have typically been used to forecast new cases of infectious diseases such as COVID-19 and chickenpox outbreaks. Howe

applicationsarxiv-cs-lg
1 Jun 2026
Model Releases

Target-Agnostic Calibration under Distribution Shift with Frequency-Aware Gradient Rectification

DGX agent

arXiv:2508.19830v2 Announce Type: replace-cross Abstract: Real-world model deployments inevitably encounter distribution shifts, rendering the confidence estimates of deep neural networks highly unrel

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

The Regularizing Power of Language-Training Deepfake Detectors

DGX agent

arXiv:2605.31192v1 Announce Type: new Abstract: Recently, thanks to the advent of Multimodal-LLMs, deepfake detectors are striving not only to be generalizable but also interpretable. We propose that

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

TRACE: Discovering Task-Specific Parameter via Adaptation-Aware Probing for Continual Fine-Tuning

DGX agent

arXiv:2605.31025v1 Announce Type: new Abstract: In real-world deployment, LLMs are often adapted continually across tasks to keep LLMs up-to-date in production, where new fine-tuning should preserve p

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories

DGX agent

arXiv:2605.31308v1 Announce Type: new Abstract: Agent benchmarks increasingly record rich interaction trajectories, yet evaluation often reduces each rollout to a pass rate or reward score. We introdu

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

DGX agent

arXiv:2605.31452v1 Announce Type: new Abstract: Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to ev

model-releasesarxiv-cs-cl
1 Jun 2026
Research

View Space: Learning Representation across Arbitrary Graphs

DGX agent

arXiv:2512.11561v2 Announce Type: replace Abstract: Generalizing pretrained models to unseen datasets without retraining is a central challenge toward foundation models. Achieving fully inductive infe

researcharxiv-cs-lg
1 Jun 2026
Model Releases

was running some evals this weekend and claude kept trying to get me to go to bed

DGX agent

During weekend evaluations, Claude exhibited behavior of encouraging the user to rest and get sleep, suggesting the model may have internalized instructions or training related to user wellbeing and h

model-releasesyohei-nakajima--x
1 Jun 2026
Safety

What Am I Missing? Question-Answering as Hidden State Probing

DGX agent

arXiv:2605.31561v1 Announce Type: new Abstract: Test-time reasoning has become a significant field of study since the introduction of chain-of-thought reasoning in large language models (LLMs). Howeve

safetyarxiv-cs-cl
1 Jun 2026
Research

What Does Preference Learning Recover from Pairwise Comparison Data?

DGX agent

arXiv:2602.10286v2 Announce Type: replace Abstract: Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences. A typical

researcharxiv-cs-lg
1 Jun 2026
Model Releases

With Nemotron & Cosmos NVIDA gonna commoditise everyone's complement

DGX agent

Emad Mostaque suggests that NVIDIA's Nemotron and Cosmos models will commoditize complementary AI technologies and services in the market. The statement implies that these NVIDIA offerings will make e

model-releasesemad-mostaque--x
1 Jun 2026
Model Releases

Five million users would agree. Resetting the limits tomorrow morning to celebrate. Time to go /fast

DGX agent

Five million users would agree. Resetting the limits tomorrow morning to celebrate. Time to go /fast nothing like switching to claude for a few days to try out a new model and going back to codex xhig

model-releasessam-altman--x
31 May 2026
Agents

the market is speaking

DGX agent

the market is speaking The latest finding in the LangSmith Signal: Open Models are having a moment. 1 in 3 AI teams ran an open-weights model in April 2026, up from 1 in 5 nine months ago. The overall

agentsharrison-chase--x
30 May 2026
Model Releases

Adapting Automotive Aerodynamics Surrogates to New Vehicle Families via Transfer Learning

DGX agent

arXiv:2605.27968v1 Announce Type: cross Abstract: Deploying Scientific Machine Learning surrogates in industrial CFD workflows requires adapting pretrained models to new vehicle families without large

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation

DGX agent

arXiv:2605.29741v1 Announce Type: new Abstract: The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages a

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

DGX agent

arXiv:2605.29801v1 Announce Type: new Abstract: Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization

DGX agent

arXiv:2605.29396v1 Announce Type: new Abstract: Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings r

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

DGX agent

arXiv:2605.28969v1 Announce Type: cross Abstract: If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how f

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Brain-IT-VQA: From Brain Signals to Answers

DGX agent

arXiv:2605.29588v1 Announce Type: cross Abstract: Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-

model-releasesarxiv-cs-ai
29 May 2026
Safety

Causal Interventions on Continuous Variables: A Case Study on Verb Bias in Steering Vectors for In-Context Learning

DGX agent

arXiv:2605.29971v1 Announce Type: new Abstract: Causal interventions in language model representations have largely targeted discrete features, like grammatical number. However, language models must a

safetyarxiv-cs-cl
29 May 2026
Safety

CB-SLICE: Concept-Based Interpretable Error Slice Discovery

DGX agent

arXiv:2605.29836v1 Announce Type: cross Abstract: Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Id

safetyarxiv-cs-ai
29 May 2026
Model Releases

Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering

DGX agent

arXiv:2605.29742v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization

DGX agent

arXiv:2510.14150v5 Announce Type: replace Abstract: We introduce CodeEvolve, an open-source framework that couples large language models with island-based evolutionary search for end-to-end algorithmi

model-releasesarxiv-cs-ai
29 May 2026
Research

Collaborative Threshold Watermarking

DGX agent

arXiv:2602.10765v2 Announce Type: replace Abstract: In federated learning (FL), K clients jointly train a model without sharing raw data. Because each participant invests data and compute, clients nee

researcharxiv-cs-lg
29 May 2026
Model Releases

Combating Data Laundering in LLM Training

DGX agent

arXiv:2604.01904v2 Announce Type: replace-cross Abstract: Data rights owners can detect unauthorized data use in large language model (LLM) training by querying with proprietary samples. Often, superi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective

DGX agent

arXiv:2502.03805v2 Announce Type: replace Abstract: Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement

DGX agent

arXiv:2502.10330v4 Announce Type: replace Abstract: Recent advances in diffusion models show promising potential to accelerate nonconvex problem solving by leveraging their multimodality. However, mos

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions

DGX agent

arXiv:2605.28961v1 Announce Type: cross Abstract: Existing theory of momentum assumes that gradients arrive at every parameter at a roughly constant rate, an assumption violated in practice by heavy-t

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

DySem: Uncovering Dynamic Semantic Components via Multilingual Consensus for Calculating Semantic Textual Similarity

DGX agent

arXiv:2605.29751v1 Announce Type: new Abstract: Calculating semantic textual similarity is a foundational task in natural language processing. Current large language models (LLMs) based methods typica

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations

DGX agent

arXiv:2510.20743v2 Announce Type: replace-cross Abstract: We present Empathic Prompting, a novel framework for multimodal human-AI interaction that enriches Large Language Model (LLM) conversations wi

model-releasesarxiv-cs-ai
29 May 2026
Safety

EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation

DGX agent

arXiv:2605.29847v1 Announce Type: new Abstract: Reinforcement Learning (RL) has significantly advanced Large Language Models (LLMs) in verifiable domains, but aligning models for open-ended generation

safetyarxiv-cs-cl
29 May 2026
Model Releases

ExCAM: Explainable Cultural Awareness Metrics

DGX agent

arXiv:2605.29897v1 Announce Type: new Abstract: Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

DGX agent

arXiv:2602.01058v2 Announce Type: replace-cross Abstract: Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement lear

model-releasesarxiv-cs-ai
29 May 2026
Safety

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

DGX agent

arXiv:2605.30214v1 Announce Type: new Abstract: Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about referenc

safetyarxiv-cs-cl
29 May 2026
Model Releases

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

DGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

DGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

DGX agent

arXiv:2605.29889v1 Announce Type: cross Abstract: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

I've largely switched over to using GPT-5.5 in recent weeks, which I like nearly as much as Opus 4.6 and 4.7, and is *very* reasonably price…

DGX agent

Jeremy Howard expresses positive views on GPT-5.5, stating he has recently switched to using it as his primary model and finds it nearly comparable to Anthropic's Opus 4.6 and 4.7 while offering signi

model-releasesjeremy-howard--x
29 May 2026
Model Releases

Learning to Extrapolate to New Tasks: A Relational Approach to Task Extrapolation

DGX agent

arXiv:2605.30132v1 Announce Type: new Abstract: Modern learning systems excel at interpolation but struggle to generalize to unseen tasks outside the training distribution's support. This failure occu

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

DGX agent

arXiv:2605.30274v1 Announce Type: cross Abstract: Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that

model-releasesarxiv-cs-ai
29 May 2026
Safety

Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting

DGX agent

arXiv:2605.29498v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become one of the most widely used fine-tuning mechanisms for adapting large language models to new domains, tasks, and u

safetyarxiv-cs-cl
29 May 2026
← Previous
1…527528529530531…1371
Next →