AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlog
91,020Total entries
1Added by human
91,019Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,767 results
Tutorials

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

DGX agent

arXiv:2602.08222v2 Announce Type: replace Abstract: As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow hi

tutorialsarxiv-cs-ai
9 Jun 2026
Model Releases

What Codex unlocks for Notion

DGX agent
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

OpenAI's Codex model enables Notion to add AI-powered capabilities to its workspace platform, allowing users to automate tasks and generate content through natural language commands. This integration

model-releasesopenai
9 Jun 2026
Safety

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO

DGX agent

arXiv:2606.06828v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. Howeve

safetyarxiv-cs-cv
8 Jun 2026
Model Releases

Aumann-SHAP: The Geometry of Counterfactual Interaction Explanations in Machine Learning

DGX agent

arXiv:2603.14014v2 Announce Type: replace Abstract: We introduce Aumann-SHAP, an interaction-aware framework that decomposes counterfactual transitions by restricting the model to a local hypercube co

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

Building super fast experiences with Gemma just got easier. Gemma 4 MTP is now officially merged into llama.cpp. Developers can now pair MTP…

DGX agent

Gemma 4 MTP (Multi-Token Prediction) has been officially integrated into llama.cpp, enabling developers to build faster AI experiences by using the model with this inference framework. This merge allo

model-releasesgeorgi-gerganov--x
8 Jun 2026
Research

CoMetaPNS: Continually Meta-learning Personalized Neural Surrogates for Cardiac Electrophysiology Simulations

DGX agent

arXiv:2606.07488v1 Announce Type: new Abstract: Personalized virtual heart simulations face challenges in model personalization and computational cost. While neural surrogates offer state-of-the-art s

researcharxiv-cs-lg
8 Jun 2026
Model Releases

Compute-Optimal Network Design for Echocardiography Myocardial Segmentation and Perfusion Quantification using Neural Scaling Laws

DGX agent

arXiv:2606.06725v1 Announce Type: cross Abstract: Myocardial perfusion quantification using contrast-enhanced ultrasound offers a bedside non-ionizing alternative to nuclear imaging modalities. Howeve

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval

DGX agent

arXiv:2506.11066v3 Announce Type: replace-cross Abstract: Code retrieval is essential in modern software development, as it boosts code reuse and accelerates debugging. However, current benchmarks pri

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

DaX: Learning General Pathology Representations Across Scales

DGX agent

arXiv:2606.06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnificatio

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Declarative Skills for AI Agents in Knowledge-Grounded Tool-Use Workflows

DGX agent

arXiv:2606.06923v1 Announce Type: new Abstract: We study orchestration mechanisms for tool-using AI agents in realistic customer-service workflows over an unstructured knowledge base. We argue that de

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

From Vision to Text: A Compact Multimodal Approach for Robust, Cross-Domain Presentation Attack Detection on ID Cards

DGX agent

arXiv:2606.06966v1 Announce Type: new Abstract: Cross-domain shifts challenge Presentation Attack Detection (PAD) on ID Cards, given the restricted data available due to privacy concerns. This work pr

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

Gemma 4 Chat Template now has preserve thinking

DGX agent

Google added an empty thinking token to the Gemma 4 chat template, which stabilizes model output by suppressing 'ghost' thought channels that may appear even when thinking is deactivated. This update

model-releasesr-localllama
8 Jun 2026
Model Releases

Learning Perspectivist Social Meaning via Demographic-Conditioned Fusion Embeddings

DGX agent

arXiv:2606.07123v1 Announce Type: new Abstract: Social meaning in language is inherently perspectival, varying across annotator backgrounds, demographics, and ideological positions. However, most NLP

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

LLM-Guided Evolution for Medical Decision Pipelines

DGX agent

arXiv:2606.07342v1 Announce Type: new Abstract: Adapting large language models (LLMs) to clinical workflows often requires costly fine-tuning or manual prompt and pipeline engineering. We study LLM-gu

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

MacArena: Benchmarking Computer Use Agents on an Online macOS Environment

DGX agent

arXiv:2606.06560v1 Announce Type: cross Abstract: Computer-use agents (CUAs) operate graphical user interfaces (GUIs) through vision and control primitives, and their capabilities have advanced rapidl

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

NotebookLM’s Gemini 3.5 upgrade adds a cloud computer and help finding sources

DGX agent

Google is rolling out 'across the board' updates to NotebookLM. The AI-powered note-taking app now uses Google's upgraded Gemini 3.5 model, which will allow it to respond with 'more accurate and relia

model-releasesthe-verge-ai
8 Jun 2026
Model Releases

On the Geometry of On-Policy Distillation

DGX agent

arXiv:2606.07082v1 Announce Type: cross Abstract: On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We ch

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese

DGX agent

arXiv:2606.07300v1 Announce Type: new Abstract: Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning. However, most large language model (LLM) research focuses on meanin

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

Pipeline parallelism in llama.cpp may be wasting your VRAM

DGX agent

Pipeline parallelism in llama.cpp distributes model layers across multiple GPUs, with each GPU holding a contiguous slice of layers . However, the Reddit post likely discusses inefficiencies in how pi

model-releasesr-localllama
8 Jun 2026
Safety

RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning

DGX agent

arXiv:2606.07006v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations,

safetyarxiv-cs-cl
8 Jun 2026
Research

Re-Centering Humans in LLM Personalization

DGX agent

arXiv:2606.06614v1 Announce Type: cross Abstract: Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data. It remains uncle

researcharxiv-cs-ai
8 Jun 2026
Model Releases

ScenicRules: An Autonomous Driving Benchmark with Multi-Objective Specifications and Abstract Scenarios

DGX agent

arXiv:2602.16073v2 Announce Type: replace-cross Abstract: Developing autonomous driving systems for complex traffic environments requires balancing multiple objectives, such as avoiding collisions, ob

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

DGX agent

arXiv:2606.06837v1 Announce Type: cross Abstract: Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

Siri AI at WWDC 2026

DGX agent

Given how badly burned anyone who took Apple's 2024 WWDC Apple Intelligence announcements at face value was, I'm holding to a strict 'I'll believe it when I see it' policy for everything they announce

model-releasessimon-willison
8 Jun 2026
Model Releases

So excited to be opening up OpenEnv to the whole community. It will now be owned by @huggingface , Meta-PyTorch, @reflection_ai , @UnslothAI…

DGX agent

So excited to be opening up OpenEnv to the whole community. It will now be owned by @huggingface , Meta-PyTorch, @reflection_ai , @UnslothAI , @modal, @PrimeIntellect , @NVIDIAAI , @mercor_ai , and @f

model-releasesclem-delangue--x
8 Jun 2026
Model Releases

Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning

DGX agent

arXiv:2606.07500v1 Announce Type: cross Abstract: Continual learning in Large Language Models (LLMs) is hindered by the plasticity-stability dilemma, where acquiring new capabilities often leads to ca

model-releasesarxiv-cs-ai
8 Jun 2026
Safety

Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry

DGX agent

arXiv:2603.26846v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical. A vital risk is intrinsic

safetyarxiv-cs-ai
8 Jun 2026
Local Ai

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

DGX agent

arXiv:2606.07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural kn

local-aiarxiv-cs-ai
8 Jun 2026
Research

The Necessity of Setting Temperature in LLM-as-a-Judge

DGX agent

arXiv:2603.28304v2 Announce Type: replace Abstract: Using large language models (LLMs) as judges for evaluating model outputs has emerged as an important paradigm for automated evaluation. However, th

researcharxiv-cs-cl
8 Jun 2026
Model Releases

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

DGX agent

arXiv:2606.06667v1 Announce Type: new Abstract: The mechanisms behind LLMs' broad over-generalization beyond training examples remain unclear. Emergent misalignment (EM) offers a striking case study:

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning

DGX agent

arXiv:2606.06915v1 Announce Type: cross Abstract: Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language model (LLM) reasoning by allocating additional compute

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

DGX agent

arXiv:2606.06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak g

model-releasesarxiv-cs-lg
8 Jun 2026
Safety

What Do People Actually Want From AI? Mapping Preference Plurality

DGX agent

arXiv:2606.06674v1 Announce Type: new Abstract: Large Language Models (LLMs) are often fine-tuned through Reinforcement Learning from Human Feedback (RLHF) to align with people's preferences and value

safetyarxiv-cs-cl
8 Jun 2026
Model Releases

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning

DGX agent

arXiv:2606.06745v1 Announce Type: new Abstract: Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is c

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs

DGX agent

arXiv:2601.12359v1 Announce Type: cross Abstract: Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Wasn't Krea 2 supposed to be released ?

DGX agent

Krea 2, Krea's first foundation image model built from scratch, was announced on May 12, 2026 , with a focus on aesthetics, style transfer, and creative control . Krea 2 became available to everyone s

model-releasesr-stablediffusion
7 Jun 2026
Model Releases

Agent-Orchestrated Adaptive RAG: A Comparative Study on Structured and Multi-Hop Retrieval

DGX agent

arXiv:2606.05658v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding their responses in external knowledge, but conventional pipeli

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Enhancing Software Engineering Through Closed-Loop Memory Optimization

DGX agent

arXiv:2606.05646v1 Announce Type: cross Abstract: Large language models (LLMs) have enabled powerful software engineering (SE) agents capable of navigating complex codebases and resolving real-world i

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Evaluation of LLMs for Mathematical Formalization in Lean

DGX agent

arXiv:2606.05632v1 Announce Type: new Abstract: Within the past few years, the ability of Large Language Models (LLMs) to generate formal mathematical proofs has improved drastically. We provide a com

model-releasesarxiv-cs-ai
6 Jun 2026
Agents

No frontier lab will own every single point on the pareto frontier around cost/latency and accuracy. Even as the pareto frontier itself adva…

DGX agent

No frontier lab will own every single point on the pareto frontier around cost/latency and accuracy. Even as the pareto frontier itself advances, there will always be points owned by open-weight model

agentsjerry-liu--x
6 Jun 2026
Model Releases

SagnacAssisted Enhanced OTDR for Distributed Acoustic Sensing: A Standardized Benchmark and Engineering Evaluation Framework

DGX agent

arXiv:2606.05754v1 Announce Type: cross Abstract: Phase-sensitive optical time-domain reflectometry (phi-OTDR) is widely used in large-scale distributed acoustic sensing (DAS) because it provides dist

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Synthetic Contrastive Reasoning for Multi-Table Q&A

DGX agent

arXiv:2606.05382v1 Announce Type: new Abstract: Multi-table question answering requires models to retrieve relevant evidence, link schemas, and perform compositional reasoning across relational tables

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems

DGX agent

arXiv:2606.05304v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on large language models are typically organized around roles, pipelines, and turn schedules, while the content that age

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

DGX agent

arXiv:2606.05806v1 Announce Type: new Abstract: Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures. We intr

model-releasesarxiv-cs-ai
6 Jun 2026
Research

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

DGX agent

arXiv:2510.05544v2 Announce Type: replace Abstract: Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and comp

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs

DGX agent

arXiv:2602.09574v2 Announce Type: replace Abstract: Tree-search decoding is an effective form of test-time scaling for large language models (LLMs), but real-world deployment often imposes a fixed per

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

Alignment Risks from Capability-Seeking RL Training

DGX agent

arXiv:2602.12124v2 Announce Type: replace-cross Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capab

safetyarxiv-cs-cl
5 Jun 2026
Model Releases

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

DGX agent

arXiv:2606.05553v1 Announce Type: new Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Exis

model-releasesarxiv-cs-cl
5 Jun 2026
← Previous
1…521522523524525…1371
Next →