AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

87,042Total entries
1Added by human
87,041Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,509 results
9 Jun 2026

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

Model ReleasesDGX agent

arXiv:2606.07963v1 Announce Type: new Abstract: Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific trigg

Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms

ResearchDGX agent

arXiv:2606.08236v1 Announce Type: cross Abstract: As large language models are increasingly deployed in high-stakes settings, there is a growing need for tools that audit not only model outputs but al

SNN-MLIR: An MLIR Dialect for Compiling Neuromorphic SNNs from NIR to Bare-Metal C

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.09213v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) are increasingly trained in a wide range of frameworks (SnnTorch, Lava, Norse, and others) each with its own model form

Sovereign AI for all.

Model ReleasesDGX agent

Cohere advocates for democratizing access to sovereign AI systems, enabling organizations and nations to develop and deploy their own AI models independently rather than relying on centralized provide

SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving

Model ReleasesDGX agent

arXiv:2606.08635v1 Announce Type: new Abstract: Prefill-decode (PD) disaggregation decouples prompt processing from token generation, but it also turns the key-value (KV) cache into a network payload.

Steganography Without Modification: Hidden Communication via LLM Seeds

SafetyDGX agent

arXiv:2606.09135v1 Announce Type: cross Abstract: We demonstrate that widely deployed Large Language Model (LLM) inference stacks harbor a steganographic channel that requires no modification to model

Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking

Model ReleasesDGX agent

arXiv:2606.07689v1 Announce Type: new Abstract: Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with r

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

Model ReleasesDGX agent

arXiv:2606.09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedde

The Mirrored Influence Hypothesis: Efficient Data Influence Estimation by Harnessing Forward Passes

ResearchDGX agent

arXiv:2402.08922v3 Announce Type: replace Abstract: Large-scale black-box models have become ubiquitous across numerous applications. Understanding the influence of individual training data sources on

There has been a lot of hand wringing on the appropriate valuation of SpaceX. Some large institutions believe SpaceX can only be valued at h…

Model ReleasesDGX agent

There has been a lot of hand wringing on the appropriate valuation of SpaceX. Some large institutions believe SpaceX can only be valued at half what the market seems to be willing to pay for it. Other

TQA-Bench: Evaluating LLMs for Multi-Table Question Answering

Model ReleasesDGX agent

arXiv:2411.19504v2 Announce Type: replace Abstract: The advance of large language models (LLMs) has unlocked great opportunities in complex multi-modal data management tasks, particularly in question

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

Model ReleasesDGX agent

arXiv:2606.09323v1 Announce Type: new Abstract: Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare d

Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks,Challenges and Baselines

Model ReleasesDGX agent

arXiv:2606.07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detectio

UniQL: Towards Dialect-Universal Benchmarking for Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.08018v1 Announce Type: new Abstract: Existing text-to-SQL benchmarks are largely centered on SQLite, making it difficult to evaluate whether models can generalize across heterogeneous SQL d

VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion

ResearchDGX agent

arXiv:2510.03244v2 Announce Type: replace-cross Abstract: Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores c

We've got 13 days to burn as much tokens as humanly possible on Claude Max plans Before they revert to API based billing 💀

Model ReleasesDGX agent

We've got 13 days to burn as much tokens as humanly possible on Claude Max plans Before they revert to API based billing 💀 Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for gen

8 Jun 2026

A Geometric Gaussian Mixture Representation of Plane Curves

Model ReleasesDGX agent

arXiv:2606.06505v1 Announce Type: cross Abstract: We introduce a user defined probabilistic polygonal representation for plane curves. Given a curve, we select vertices on the curve and connect consec

Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle

Model ReleasesDGX agent

arXiv:2606.07462v1 Announce Type: new Abstract: As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing

ResearchDGX agent

arXiv:2606.06924v1 Announce Type: new Abstract: Existing LLM routing methods typically treat a model's single response to a query as its capability label for training routers. However, because LLM gen

Generative Molecular Morphing for Flexible-Size Design via Unbalanced Optimal Transport

ResearchDGX agent

arXiv:2606.07239v1 Announce Type: new Abstract: The success of generative molecular design hinges on a model's steerability toward high-reward samples. Because many molecular properties are intrinsica

Hierarchical Certified Semantic Commitment for Byzantine-Resilient LLM-Agent Collaboration

Model ReleasesDGX agent

arXiv:2606.07316v1 Announce Type: cross Abstract: Byzantine collaboration among large-language-model agents requires a finality-control primitive: given delivered stochastic, structured natural-langua

Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

Model ReleasesDGX agent

arXiv:2606.07033v1 Announce Type: new Abstract: Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including catego

Improving Cross-Lingual Factual Recall via Consistency-Driven Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.06586v1 Announce Type: new Abstract: Large language models (LLMs) trained predominantly on English data encode substantial world knowledge, yet often fail to express it reliably in other la

It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

Model ReleasesDGX agent

arXiv:2512.23128v2 Announce Type: replace-cross Abstract: Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their r

MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

Model ReleasesDGX agent

arXiv:2606.07512v1 Announce Type: cross Abstract: Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and

Online Pandora's Box for Contextual LLM Cascading

SafetyDGX agent

arXiv:2606.07392v1 Announce Type: new Abstract: Motivated by Large Language Model (LLM) cascading, we propose an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs. I

PromptPrint: Behavioral Biometrics Through Natural Language Prompting in LLMs

Model ReleasesDGX agent

arXiv:2606.06755v1 Announce Type: new Abstract: Authorship attribution research has traditionally focused on long-form, expressive texts; however, interactions with large language models (LLMs) are ty

Real-Time AttentionBender: Granular Interactive Network Bending of Video Diffusion Transformers

ResearchDGX agent

arXiv:2606.06497v1 Announce Type: cross Abstract: Generative video models have achieved remarkable visual fidelity, yet their prompt-only interface offers thin creative agency and obscures the model's

REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference

Model ReleasesDGX agent

arXiv:2606.07141v1 Announce Type: cross Abstract: Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, and data owne

RETROSPECT: RETROsynthesis via Sequential Prediction, and Chemically Transformed-ranking

Model ReleasesDGX agent

arXiv:2606.07181v1 Announce Type: cross Abstract: Single-step retrosynthesis needs both accurate first-ranked suggestions and candidate lists that are rich enough for downstream selection. We study th

Scalable Joint Resource Allocation for SLO-Constrained LLM Inference in Heterogeneous GPU Clouds

HardwareDGX agent

arXiv:2604.07472v2 Announce Type: replace Abstract: Serving large language model (LLM) inference in cloud environments requires jointly optimizing model selection, GPU provisioning, parallelism config

Seeing Without Exposing: Adaptive Privacy Control for Open-World, Context-Hungry MLLMs

Model ReleasesDGX agent

arXiv:2606.07175v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have raised new privacy challenges. On the data side, user-provided inputs often include unpredictable sensitiv

SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices

Model ReleasesDGX agent

arXiv:2606.07098v1 Announce Type: new Abstract: We present SigmaScale, a method for learning auxiliary scaling matrices S to aid truncated Singular Value Decomposition (SVD) based Large Language Model

STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation

Model ReleasesDGX agent

arXiv:2606.07036v1 Announce Type: cross Abstract: Synthetic histopathology image generation addresses critical challenges in computational pathology, including patient privacy and the growing need for

TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment

Model ReleasesDGX agent

arXiv:2606.07451v1 Announce Type: cross Abstract: Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and te

VeriDrive: Verifiable Counterfactual Supervision for Cost-Efficient Vision-Language Planning

Model ReleasesDGX agent

arXiv:2606.07338v1 Announce Type: new Abstract: Vision-language driving models increasingly use reasoning supervision to bridge perception, prediction, and planning, but existing driving rationales ar

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing

Model ReleasesDGX agent

arXiv:2606.07171v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) enable flexible instruction-driven image editing, but privacy risks arise when user images expose diverse and u

6 Jun 2026

Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillatio

Model ReleasesDGX agent

arXiv:2606.05682v1 Announce Type: new Abstract: Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost c

Evaluating Agentic Configuration Repair for Computer Networks

Model ReleasesDGX agent

arXiv:2606.06212v1 Announce Type: new Abstract: Misconfigurations in computer networks remain a major source of critical Internet outages. Research is turning to Large Language Models (LLMs) to automa

Geographic Bias and Diversity in AI Evaluation

Model ReleasesDGX agent

arXiv:2606.05187v1 Announce Type: cross Abstract: Among the many challenges hindering the responsible development and deployment of AI, arguably none has faced more intense scrutiny than bias in its v

GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

Model ReleasesDGX agent

arXiv:2606.05566v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attack

Multilingual Fine-Tuning via Localized Gradient Conflict Resolution

Model ReleasesDGX agent

arXiv:2606.05613v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) has established cross-lingual versatility as a defining feature of modern systems. However, fine-tun

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

ResearchDGX agent

arXiv:2606.05555v1 Announce Type: cross Abstract: Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge. While recent advances in model-based RL achieve strong

SentinelBench: A Benchmark for Long-Running Monitoring Agents

Model ReleasesDGX agent

arXiv:2606.05342v1 Announce Type: new Abstract: AI agents are increasingly asked to carry out work that spans minutes, hours, or longer. Yet the default model of agent behavior is continuous action: i

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability

Model ReleasesDGX agent

arXiv:2606.06333v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their formulation assigns each latent featur

TokenMizer: Graph-Structured Session Memory for Long-Horizon LLM Context Management

Model ReleasesDGX agent

arXiv:2606.06337v1 Announce Type: new Abstract: Large language model (LLM) deployments for long-horizon tasks face a fundamental constraint: context windows are finite while productive work sessions a

ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents

Model ReleasesDGX agent

arXiv:2606.06284v1 Announce Type: new Abstract: Large language model agents increasingly rely on external tools, but larger tool menus can reduce reliability and efficiency by increasing wrong-tool ca

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents

Model ReleasesDGX agent

arXiv:2606.06453v1 Announce Type: new Abstract: Sparse attention is becoming increasingly important for serving large language models (LLMs) as generation lengths continue to grow. However, deploying

When Should Memory Stay Silent: Measuring Memory-Use Boundaries in Memory-Augmented Conversational Agents

Model ReleasesDGX agent

arXiv:2606.06055v1 Announce Type: new Abstract: Long-term memory enables language model agents to support personalized interactions, but it remains unclear when available memories warrant integration

5 Jun 2026

Arena AI Agentic User Benchmark Ranking

Model ReleasesDGX agent

Arena AI's agentic benchmark ranks AI models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability. The leaderboard

At least until (if?) rapid improvement stops, it seems less likely someone is going to catch the Big Three AI Labs. Microsoft and Meta relea…

Model ReleasesDGX agent

At least until (if?) rapid improvement stops, it seems less likely someone is going to catch the Big Three AI Labs. Microsoft and Meta released their models, which were fine, but not frontier. SpaceX

CLEAR: Cognition and Latent Evaluation for Adaptive Routing in End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.06219v1 Announce Type: new Abstract: End-to-end autonomous driving models often struggle to balance multi-modal maneuver generation with real-time inference constraints. While diffusion mod

CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement

Model ReleasesDGX agent

arXiv:2606.05793v1 Announce Type: new Abstract: While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conver

DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections

Model ReleasesDGX agent

arXiv:2508.15851v2 Announce Type: replace Abstract: Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information

Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based Learning

TutorialsDGX agent

arXiv:2606.05471v1 Announce Type: new Abstract: Learning semantics is essential for deep learning models to be interpretable and better aligned with human reasoning. Concept-based models approach this

From Self to Other: Evaluating Demographic Perspective-Taking in LLM Hate Speech Annotation

Model ReleasesDGX agent

arXiv:2606.06266v1 Announce Type: new Abstract: Hate speech detection is inherently subjective: people from different demographic groups perceive the same content very differently. Collecting enough a

Generic Triple-Latent Compression with Gated Associative Retrieval

Model ReleasesDGX agent

arXiv:2606.05175v1 Announce Type: new Abstract: We study generic triple-latent sequence models that maintain a running token state and compressed pair-memory pathway to capture higher-order token inte

https://ollama.com/library/gemma4/tags

Local AiDGX agent

Gemma4 is a language model available through Ollama's model library with multiple tagged versions for different use cases and configurations. The Ollama platform enables users to run open-source large

IA-RAG: Interval-Algebra-Driven Temporal Reasoning for Dynamic Knowledge Retrieval

Model ReleasesDGX agent

arXiv:2606.06044v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has shown strong effectiveness in grounding Large Language Models (LLMs) with external knowledge. However, existing

Improving Heart-Focused Medical Question Answering in LLMs via Variance-Aware Rubric Rewards with GRPO

Local AiDGX agent

arXiv:2606.05174v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong promise in healthcare applications. Yet deploying general-purpose models in real-world settings remains d

← Previous
1…343344345346347…1042
Next →