AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,082 results
6 May 2026

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use

Model ReleasesDGX agent

arXiv:2605.02964v1 Announce Type: new Abstract: Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomou

Self-Mined Hardness for Safety Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.03226v1 Announce Type: new Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's diff

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It

Model ReleasesDGX agent

arXiv:2605.03258v1 Announce Type: cross Abstract: Large language models often fail at simple counting tasks, even when the items to count are explicitly present in the prompt. We investigate whether t

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

What's new in IAM: Security, governance, and runtime defense

Model ReleasesDGX agent

The AI era demands a fundamental shift in security, and that includes identity and access management (IAM). Traditional controls simply aren’t built for autonomous AI agents that interact with sensiti

When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

Model ReleasesDGX agent

arXiv:2605.03096v1 Announce Type: cross Abstract: In classification tasks, models may rely on confounding variables to achieve strong in-distribution performance, capturing spurious features that fail

5 May 2026

A Light Weight Multi-Features-View Convolution Neural Network For Plant Disease Identification

Model ReleasesDGX agent

arXiv:2605.00903v1 Announce Type: new Abstract: Agriculture is a key sector of the economies of developing countries. It serves as a primary source of income and employment for rural populations. Howe

Adaptive Texture-aware Masking for Self-Supervised Learning in 3D Dental CBCT Analysis

Model ReleasesDGX agent

arXiv:2605.01741v1 Announce Type: new Abstract: Cone Beam Computed Tomography (CBCT) is pivotal for 3D diagnostic imaging in dentistry. However, the development of robust AI models for volumetric anal

Anon: Extrapolating Optimizer Adaptivity Across the Real Spectrum

ResearchDGX agent

arXiv:2605.02317v1 Announce Type: cross Abstract: Adaptive optimizers such as Adam have achieved great success in training large-scale models like large language models and diffusion models. However,

ARIS: Agentic and Relationship Intelligence System for Social Robots

Model ReleasesDGX agent

arXiv:2605.00943v1 Announce Type: new Abstract: Foundational models have advanced social robotics, enabling richer perception and communicative interaction with users. However, current systems still s

Compute Optimal Tokenization

Model ReleasesDGX agent

arXiv:2605.01188v1 Announce Type: new Abstract: Scaling laws enable the optimal selection of data amount and language model size, yet the impact of the data unit, the token, on this relationship remai

Embedding-based In-Context Prompt Training for Enhancing LLMs as Text Encoders

Model ReleasesDGX agent

arXiv:2605.01372v1 Announce Type: new Abstract: Large language models (LLMs) have been widely explored for embedding generation. While recent studies show that in-context learning (ICL) effectively en

Five must-have guides to move agents into production with Gemini Enterprise Agent Platform

Model ReleasesDGX agent

Building AI agents that work well in a demo is one thing, but running them in production requires serious infrastructure. At Google Cloud Next '26, we introduced Gemini Enterprise Agent Platform to he

Momentum-Anchored Multi-Scale Fusion Model for Long-Tailed Chest X-Ray Classification

SafetyDGX agent

arXiv:2605.02292v1 Announce Type: new Abstract: Chest X-ray classification suffers from severe class imbalance where gradient updates bias toward majority classes, causing feature drift and poor perfo

Multi-fidelity surrogates for mechanics of composites: from co-kriging to multi-fidelity neural networks

Model ReleasesDGX agent

arXiv:2605.02871v1 Announce Type: cross Abstract: Composite materials exhibit strongly hierarchical and anisotropic properties governed by coupled mechanisms spanning constituents, plies, laminates, s

Multimodal Confidence Modeling in Audio-Visual Quality Assessment

ApplicationsDGX agent

arXiv:2605.01219v1 Announce Type: cross Abstract: Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortion

On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility

Model ReleasesDGX agent

arXiv:2605.01357v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at long-context understanding but exhibit significant limitations in long-form generation. Existing studies primarily

PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction

Model ReleasesDGX agent

arXiv:2605.01945v1 Announce Type: new Abstract: Tandem mass spectrometry provides a high-throughput framework for identifying and quantifying proteins in complex biological samples. In computational p

Robust Parameter Learning for Uncertain MDPs

Model ReleasesDGX agent

arXiv:2605.01339v1 Announce Type: new Abstract: Learning-based approaches to verifying unknown Markov decision processes (MDPs) often employ uncertain MDPs. These models use, for example, confidence i

SF20K Competition 2025: Summary and findings

Model ReleasesDGX agent

arXiv:2605.01496v1 Announce Type: new Abstract: This report presents the results and findings of the first edition of the Short-Films 20K (SF20K) Competition, held in conjunction with the SLoMO Worksh

Standing on the Shoulders of Giants: Stabilized Knowledge Distillation for Cross--Language Code Clone Detection

Model ReleasesDGX agent

arXiv:2605.02860v1 Announce Type: cross Abstract: Cross-language code clone detection (X-CCD) is challenging because semantically equivalent programs written in different languages often share little

Subquadratic launches with $29M to bring 12M-token context windows to AI

Model ReleasesDGX agent

Subquadratic, a company developing a novel generative artificial intelligence model, launched today with 29 million in seed funding. The new large language model, dubbed SubQ, uses what the company ca

The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

Model ReleasesDGX agent

arXiv:2605.02398v1 Announce Type: cross Abstract: As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability -- knowing what they do not kn

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling

SafetyDGX agent

arXiv:2605.02427v1 Announce Type: cross Abstract: A recurring pattern in 'reasoning without training' is that base LLMs already assign non-trivial probability mass to correct multi-step solutions; the

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis

Model ReleasesDGX agent

arXiv:2605.00826v1 Announce Type: cross Abstract: Text-to-video retrieval enables users to find relevant video content using natural language queries, a task that has grown increasingly important with

4 May 2026

Alethia: A Foundational Encoder for Voice Deepfakes

Model ReleasesDGX agent

arXiv:2605.00251v1 Announce Type: cross Abstract: Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, dow

Caracal: Causal Architecture via Spectral Mixing

Model ReleasesDGX agent

arXiv:2605.00292v1 Announce Type: new Abstract: The scalability of Large Language Models to long sequences is hindered by the quadratic cost of attention and the limitations of positional encodings. T

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models

ResearchDGX agent

arXiv:2605.00321v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies often fail under distribution shift, suggesting that decisions may depend on spurious visual correlations rather t

Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game

Model ReleasesDGX agent

arXiv:2605.00677v1 Announce Type: new Abstract: While Large Language Models have achieved notable success on formal mathematics benchmarks such as MiniF2F, it remains unclear whether these results ste

From Prediction to Practice: A Task-Aware Evaluation Framework for Blood Glucose Forecasting

Model ReleasesDGX agent

arXiv:2605.00645v1 Announce Type: new Abstract: Clinical time-series forecasting is increasingly studied for decision support, yet standard aggregate metrics can obscure whether a model is actually us

Generative Modeling under Non-Monotone MAR Missingness via Approximate Wasserstein Gradient Flows

ResearchDGX agent

arXiv:2604.04567v2 Announce Type: replace-cross Abstract: The prevalence of missing values in data science poses a substantial risk to any further analyses. Despite a wealth of research, principled no

How I built a free, local AI powerhouse in 10 days (Ollama + Gemma 4 + Claude Cowork 3P + Browserless)

Model ReleasesDGX agent

This post documents a 10-day project to build a local AI system using open-source tools and models, specifically combining Ollama (a local LLM framework), Gemma 4 (a language model), Claude Cowork 3P,

Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation

SafetyDGX agent

arXiv:2605.00393v1 Announce Type: new Abstract: Reinforcement learning (RL) in large environments often suffers from severe computational bottlenecks, as conventional regret minimization algorithms re

The foundation of AI scalability: one team, one platform, one operating model

IndustryDGX agent

This Databricks blog post discusses how organizations can achieve AI scalability through unified infrastructure and organizational alignment, emphasizing the importance of consolidating teams, platfor

3 May 2026

Qwen3.6 vs gpt-oss:120b on Apple Silicon — three Qwen variants benchmarked, plus what works and where it does not

Model ReleasesDGX agent

This post benchmarks three Qwen3.6 model variants against gpt-oss:120b when running on Apple Silicon hardware, evaluating their performance characteristics and practical usability. It documents both t

1 May 2026

ActiNet: An Open-Source Tool for Activity Intensity Classification of Wrist-Worn Accelerometry Using Self-Supervised Deep Learning

ResearchDGX agent

arXiv:2510.01712v2 Announce Type: replace Abstract: The use of accurate and reliable open-source human activity recognition (HAR) models on passively collected wrist-accelerometer data is essential in

Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs

Model ReleasesDGX agent

arXiv:2604.27006v1 Announce Type: cross Abstract: Context: Study screening in systematic literature reviews is costly, inconsistency-prone, and risk-asymmetric, since false negatives can compromise va

Characterizing the Consistency of the Emergent Misalignment Persona

Model ReleasesDGX agent

arXiv:2604.28082v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) on narrowly misaligned data generalizes to broadly misaligned behavior, a phenomenon termed emergent misalignme

DeepWeightFlow: Re-Basined Flow Matching for Generating Neural Network Weights

ResearchDGX agent

arXiv:2601.05052v2 Announce Type: replace Abstract: Building efficient and effective generative models for neural network weights has been a research focus of significant interest that faces challenge

Design Structure Matrix Modularization with Large Language Models

SafetyDGX agent

arXiv:2604.28018v1 Announce Type: cross Abstract: Design Structure Matrix (DSM) modularization, the task of partitioning system elements into cohesive modules, is a fundamental combinatorial challenge

Do What I Say: A Spoken Prompt Dataset for Instruction-Following

Model ReleasesDGX agent

arXiv:2603.09881v2 Announce Type: replace Abstract: Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompt

Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents

Local AiDGX agent

arXiv:2604.27143v1 Announce Type: cross Abstract: Recent research has demonstrated the potential of Large Language Models (LLMs) for autonomous penetration testing, particularly when using cloud-based

Fitting Horn DL Ontologies to ABox and Query Examples: A Tale of Simulation Quantifiers and Finite Models

ResearchDGX agent

arXiv:2604.26976v1 Announce Type: cross Abstract: We study the problem of fitting a description logic (DL) ontology to a given set of positive and negative examples that take the form of an ABox and a

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Model ReleasesDGX agent

arXiv:2604.27393v1 Announce Type: new Abstract: Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming inter

Modeling Spatial Extremal Dependence of Precipitation Using Distributional Neural Networks

ResearchDGX agent

arXiv:2407.08668v3 Announce Type: replace-cross Abstract: In this work, we propose a simulation-based estimation approach using generative neural networks to determine dependencies of precipitation ma

Musk v. Altman week 1: Elon Musk says he was duped, warns AI could kill us all, and admits that xAI distills OpenAI’s models

ResearchDGX agent

In the first week of the landmark trial between Elon Musk and OpenAI, Musk took the stand in a crisp black suit and tie and argued that OpenAI CEO Sam Altman and president Greg Brockman had deceived h

Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading

ResearchDGX agent

arXiv:2604.27637v1 Announce Type: new Abstract: Current Large Language Model (LLM) evaluation frameworks utilize the same static prompt template across all models under evaluation. This differs from t

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs

Model ReleasesDGX agent

arXiv:2604.27401v1 Announce Type: new Abstract: Perturbation probing generates task-specific causal hypotheses for FFN neurons in large language models using two forward passes per prompt and no backp

PVeRA: Probabilistic Vector-Based Random Matrix Adaptation

Model ReleasesDGX agent

arXiv:2512.07703v2 Announce Type: replace Abstract: Large foundation models have emerged in the last years and are pushing performance boundaries for a variety of tasks. Training or even finetuning su

RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension

Model ReleasesDGX agent

arXiv:2601.14289v2 Announce Type: replace-cross Abstract: Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables

RuC: HDL-Agnostic Rule Completion Benchmark Generation

Model ReleasesDGX agent

arXiv:2604.27780v1 Announce Type: cross Abstract: Large Language Models (LLMs) have rapidly improved in performance across code-related tasks, making their integration into Register Transfer Level (RT

SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images

Model ReleasesDGX agent

arXiv:2604.28039v1 Announce Type: new Abstract: Spectra are a prevalent yet highly information-dense form of scientific imagery, presenting substantial challenges to multimodal large language models (

30 Apr 2026

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio

Model ReleasesDGX agent

arXiv:2409.06624v4 Announce Type: replace-cross Abstract: Large Language Models (LLM) often need to be Continual Pre-Trained (CPT) to obtain unfamiliar language skills or adapt to new domains. The hug

A Systematic Comparison of Prompting and Multi-Agent Methods for LLM-based Stance Detection

Model ReleasesDGX agent

arXiv:2604.26319v1 Announce Type: new Abstract: Stance detection identifies the attitude of a text author toward a given target. Recent studies have explored various LLM-based strategies for this task

AWS Generative AI Model Agility Solution: A comprehensive guide to migrating LLMs for generative AI production

TutorialsDGX agent

In this post, we introduce a systematic framework for LLM migration or upgrade in generative AI production, encompassing essential tools, methodologies, and best practices. The framework facilitates t

> be me > 'the internet is polluted by ai slop, we need low-background tokens' > 'wouldnt it be cool if we could time travel and see what ou…

Model ReleasesDGX agent

> be me > 'the internet is polluted by ai slop, we need low-background tokens' > 'wouldnt it be cool if we could time travel and see what our ancestors 100 years ago would say to us' > all the existin

Bootstrapping Sign Language Annotations with Sign Language Models

ResearchDGX agent

AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain professional interpreters and 100s of hours of d

Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making

ResearchDGX agent

arXiv:2604.26169v1 Announce Type: new Abstract: Treatment allocation under budget constraints is a central challenge in digital advertising: advertisers must decide which users to show ads to while sp

CoQuant: Joint Weight-Activation Subspace Projection for Mixed-Precision LLMs

Model ReleasesDGX agent

arXiv:2604.26378v1 Announce Type: new Abstract: Post-training quantization (PTQ) has become an important technique for reducing the inference cost of Large Language Models (LLMs). While recent mixed-p

Emergent Coordination in Multi-Agent Language Models

Local AiDGX agent

arXiv:2510.05174v4 Announce Type: replace-cross Abstract: When are multi-agent LLM systems merely a collection of individual agents versus an integrated collective with higher-order structure? We intr

MoRFI: Monotonic Sparse Autoencoder Feature Identification

Model ReleasesDGX agent

arXiv:2604.26866v1 Announce Type: new Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of

← Previous
1…286287288289290…1035
Next →