AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

87,814Total entries
1Added by human
87,813Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,138 results
8 Jun 2026

LLM-Guided Evolution for Medical Decision Pipelines

Model ReleasesDGX agent

arXiv:2606.07342v1 Announce Type: new Abstract: Adapting large language models (LLMs) to clinical workflows often requires costly fine-tuning or manual prompt and pipeline engineering. We study LLM-gu

MacArena: Benchmarking Computer Use Agents on an Online macOS Environment

Model ReleasesDGX agent

arXiv:2606.06560v1 Announce Type: cross Abstract: Computer-use agents (CUAs) operate graphical user interfaces (GUIs) through vision and control primitives, and their capabilities have advanced rapidl

NotebookLM’s Gemini 3.5 upgrade adds a cloud computer and help finding sources

Model ReleasesDGX agent

Google is rolling out 'across the board' updates to NotebookLM. The AI-powered note-taking app now uses Google's upgraded Gemini 3.5 model, which will allow it to respond with 'more accurate and relia

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

On the Geometry of On-Policy Distillation

Model ReleasesDGX agent

arXiv:2606.07082v1 Announce Type: cross Abstract: On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We ch

Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese

Model ReleasesDGX agent

arXiv:2606.07300v1 Announce Type: new Abstract: Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning. However, most large language model (LLM) research focuses on meanin

Pipeline parallelism in llama.cpp may be wasting your VRAM

Model ReleasesDGX agent

Pipeline parallelism in llama.cpp distributes model layers across multiple GPUs, with each GPU holding a contiguous slice of layers . However, the Reddit post likely discusses inefficiencies in how pi

RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning

SafetyDGX agent

arXiv:2606.07006v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations,

Re-Centering Humans in LLM Personalization

ResearchDGX agent

arXiv:2606.06614v1 Announce Type: cross Abstract: Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data. It remains uncle

ScenicRules: An Autonomous Driving Benchmark with Multi-Objective Specifications and Abstract Scenarios

Model ReleasesDGX agent

arXiv:2602.16073v2 Announce Type: replace-cross Abstract: Developing autonomous driving systems for complex traffic environments requires balancing multiple objectives, such as avoiding collisions, ob

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

Model ReleasesDGX agent

arXiv:2606.06837v1 Announce Type: cross Abstract: Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus

Siri AI at WWDC 2026

Model ReleasesDGX agent

Given how badly burned anyone who took Apple's 2024 WWDC Apple Intelligence announcements at face value was, I'm holding to a strict 'I'll believe it when I see it' policy for everything they announce

So excited to be opening up OpenEnv to the whole community. It will now be owned by @huggingface , Meta-PyTorch, @reflection_ai , @UnslothAI…

Model ReleasesDGX agent

So excited to be opening up OpenEnv to the whole community. It will now be owned by @huggingface , Meta-PyTorch, @reflection_ai , @UnslothAI , @modal, @PrimeIntellect , @NVIDIAAI , @mercor_ai , and @f

Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning

Model ReleasesDGX agent

arXiv:2606.07500v1 Announce Type: cross Abstract: Continual learning in Large Language Models (LLMs) is hindered by the plasticity-stability dilemma, where acquiring new capabilities often leads to ca

Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry

SafetyDGX agent

arXiv:2603.26846v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical. A vital risk is intrinsic

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

Local AiDGX agent

arXiv:2606.07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural kn

The Necessity of Setting Temperature in LLM-as-a-Judge

ResearchDGX agent

arXiv:2603.28304v2 Announce Type: replace Abstract: Using large language models (LLMs) as judges for evaluating model outputs has emerged as an important paradigm for automated evaluation. However, th

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

Model ReleasesDGX agent

arXiv:2606.06667v1 Announce Type: new Abstract: The mechanisms behind LLMs' broad over-generalization beyond training examples remain unclear. Emergent misalignment (EM) offers a striking case study:

ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning

Model ReleasesDGX agent

arXiv:2606.06915v1 Announce Type: cross Abstract: Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language model (LLM) reasoning by allocating additional compute

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak g

What Do People Actually Want From AI? Mapping Preference Plurality

SafetyDGX agent

arXiv:2606.06674v1 Announce Type: new Abstract: Large Language Models (LLMs) are often fine-tuned through Reinforcement Learning from Human Feedback (RLHF) to align with people's preferences and value

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning

Model ReleasesDGX agent

arXiv:2606.06745v1 Announce Type: new Abstract: Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is c

Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs

Model ReleasesDGX agent

arXiv:2601.12359v1 Announce Type: cross Abstract: Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such

7 Jun 2026

Wasn't Krea 2 supposed to be released ?

Model ReleasesDGX agent

Krea 2, Krea's first foundation image model built from scratch, was announced on May 12, 2026 , with a focus on aesthetics, style transfer, and creative control . Krea 2 became available to everyone s

6 Jun 2026

Agent-Orchestrated Adaptive RAG: A Comparative Study on Structured and Multi-Hop Retrieval

Model ReleasesDGX agent

arXiv:2606.05658v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding their responses in external knowledge, but conventional pipeli

Enhancing Software Engineering Through Closed-Loop Memory Optimization

Model ReleasesDGX agent

arXiv:2606.05646v1 Announce Type: cross Abstract: Large language models (LLMs) have enabled powerful software engineering (SE) agents capable of navigating complex codebases and resolving real-world i

Evaluation of LLMs for Mathematical Formalization in Lean

Model ReleasesDGX agent

arXiv:2606.05632v1 Announce Type: new Abstract: Within the past few years, the ability of Large Language Models (LLMs) to generate formal mathematical proofs has improved drastically. We provide a com

No frontier lab will own every single point on the pareto frontier around cost/latency and accuracy. Even as the pareto frontier itself adva…

AgentsDGX agent

No frontier lab will own every single point on the pareto frontier around cost/latency and accuracy. Even as the pareto frontier itself advances, there will always be points owned by open-weight model

SagnacAssisted Enhanced OTDR for Distributed Acoustic Sensing: A Standardized Benchmark and Engineering Evaluation Framework

Model ReleasesDGX agent

arXiv:2606.05754v1 Announce Type: cross Abstract: Phase-sensitive optical time-domain reflectometry (phi-OTDR) is widely used in large-scale distributed acoustic sensing (DAS) because it provides dist

Synthetic Contrastive Reasoning for Multi-Table Q&A

Model ReleasesDGX agent

arXiv:2606.05382v1 Announce Type: new Abstract: Multi-table question answering requires models to retrieve relevant evidence, link schemas, and perform compositional reasoning across relational tables

What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.05304v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on large language models are typically organized around roles, pipelines, and turn schedules, while the content that age

When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

Model ReleasesDGX agent

arXiv:2606.05806v1 Announce Type: new Abstract: Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures. We intr

5 Jun 2026

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

ResearchDGX agent

arXiv:2510.05544v2 Announce Type: replace Abstract: Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and comp

Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs

Model ReleasesDGX agent

arXiv:2602.09574v2 Announce Type: replace Abstract: Tree-search decoding is an effective form of test-time scaling for large language models (LLMs), but real-world deployment often imposes a fixed per

Alignment Risks from Capability-Seeking RL Training

SafetyDGX agent

arXiv:2602.12124v2 Announce Type: replace-cross Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capab

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

Model ReleasesDGX agent

arXiv:2606.05553v1 Announce Type: new Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Exis

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement

Model ReleasesDGX agent

arXiv:2606.05920v1 Announce Type: cross Abstract: Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output. However, real web development is different. Us

CoMoL: Efficient Mixture of LoRA Experts via Dynamic Core Space Merging

Model ReleasesDGX agent

arXiv:2603.00573v2 Announce Type: replace Abstract: Large language models (LLMs) achieve remarkable performance on diverse downstream and domain-specific tasks via parameter-efficient fine-tuning (PEF

Contextualized Prompting For Stance Detection On Social Media

Model ReleasesDGX agent

arXiv:2606.06022v1 Announce Type: new Abstract: Stance detection on social media is challenging due to short, noisy, and context-dependent language. While large language models (LLMs) show zero-shot g

Deep Learning-based 3D Oral Cavity Reconstruction Using 2D Intraoral Images

ResearchDGX agent

arXiv:2606.05998v1 Announce Type: new Abstract: Oral 3D modelling is one of the most essential stages in dentistry, and many different approaches, such as impression taking and intraoral scanning, are

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments

Model ReleasesDGX agent

arXiv:2606.06217v1 Announce Type: new Abstract: When a disaster unfolds, responders must answer not only what is happening, but also why it is happening, what will happen next, and what to do now, oft

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

Model ReleasesDGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

Facial-R1: Aligning Reasoning and Recognition for Facial Emotion Analysis

Model ReleasesDGX agent

arXiv:2511.10254v2 Announce Type: replace Abstract: Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrat

Here’s this week’s shipping recap 👇 — Nano Banana 2 & Nano Banana Pro are now GA and available via the Gemini Enterprise Agent Platform, Ge…

Model ReleasesDGX agent

Here’s this week’s shipping recap 👇 — Nano Banana 2 & Nano Banana Pro are now GA and available via the Gemini Enterprise Agent Platform, Gemini API, and in @GoogleAIStudio —Co-Scientist, our new multi

HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

Model ReleasesDGX agent

arXiv:2601.02730v3 Announce Type: replace Abstract: Visual localization on standard-definition (SD) maps has emerged as a promising low-cost and scalable solution for autonomous driving. However, exis

If Claude is good enough for Nobel Prize winners it is good enough for you https://arxiv.org/abs/2606.03300

Model ReleasesDGX agent

This post appears to reference Claude AI's capabilities and performance, likely highlighting how the model has been used or endorsed by notable researchers or Nobel Prize winners to establish credibil

KV-Control: Parameter-Efficient K/V Injection for Trajectory-Controlled Text-to-Motion

Model ReleasesDGX agent

arXiv:2606.05624v1 Announce Type: new Abstract: Text-conditioned 3D human motion models now synthesize plausible motions from prompts, but practical animation and embodied-agent workflows rarely stop

Latent Reasoning with Normalizing Flows

SafetyDGX agent

arXiv:2606.06447v1 Announce Type: new Abstract: Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. H

LiAuto-GeoX: Efficient Grounded Driving Transformer

Model ReleasesDGX agent

arXiv:2606.05774v1 Announce Type: new Abstract: Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for auton

Multilingual Detection of Alzheimer's Disease from Speech: A Cross-Linguistic Transfer Learning Approach

ResearchDGX agent

arXiv:2606.05545v1 Announce Type: new Abstract: The development of multilingual Alzheimer's Disease Dementia (AD) detection models presents significant challenges due to the resource-intensive and tim

RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing

Model ReleasesDGX agent

arXiv:2605.25956v2 Announce Type: replace Abstract: Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual re

RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

Model ReleasesDGX agent

arXiv:2606.06027v1 Announce Type: cross Abstract: Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made i

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges

Model ReleasesDGX agent

arXiv:2606.05384v1 Announce Type: cross Abstract: LLM-as-judge evaluation is widely used in benchmarking pipelines, where model outputs are compared and ranked using automated evaluators. These pipeli

Symb-xMIL: Symbolic Explanations for Multiple Instance Learning in Digital Pathology

SafetyDGX agent

arXiv:2606.06224v1 Announce Type: new Abstract: Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology. Existing methods primari

TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization

Model ReleasesDGX agent

arXiv:2606.05859v1 Announce Type: new Abstract: Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive rea

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

Model ReleasesDGX agent

arXiv:2606.05665v1 Announce Type: new Abstract: Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence w

VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes

Model ReleasesDGX agent

arXiv:2606.06074v1 Announce Type: new Abstract: We introduce VZCrash, the largest publicly available dataset of real-world vehicle collision data featuring Inertial Measurement Unit (IMU) telemetry. T

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Local AiDGX agent

arXiv:2606.06113v1 Announce Type: new Abstract: Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Di

4 Jun 2026

A Geometric View of Counterfactual Behavior: Interaction of Boundary Proximity and Local Support

Local AiDGX agent

arXiv:2606.04209v1 Announce Type: new Abstract: Counterfactual explanations seek small, semantically meaningful changes to an input that alter a model's prediction, and are widely used to interpret an

A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References

Model ReleasesDGX agent

arXiv:2508.14623v2 Announce Type: replace-cross Abstract: This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objectiv

AgentJet: A Flexible Swarm Training Framework for Agentic Reinforcement Learning

HardwareDGX agent

arXiv:2606.04484v1 Announce Type: new Abstract: We present AgentJet, a distributed swarm training framework for large language model (LLM) agent reinforcement learning. Unlike centralized frameworks t

← Previous
1…398399400401402…1053
Next →