AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,603 results
28 May 2026

RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction

Model ReleasesDGX agent

arXiv:2602.13748v2 Announce Type: replace Abstract: Multimedia Event Extraction (MEE) aims to identify events and their arguments from documents that contain both text and images. It requires groundin

Robust Moment-Based Estimation via Spectral Gradient Reweighting

Model ReleasesDGX agent

arXiv:2605.27718v1 Announce Type: cross Abstract: Moment-based estimation is a theoretically attractive approach to parametric inference, especially when likelihood-based estimation is unavailable, mi

RW-TTT: Batched Serving for Request-Owned Test-Time Training State

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.28053v1 Announce Type: new Abstract: Test-time training (TTT) adapts an LLM during generation by reading and updating request-owned state, such as fast weights, low-rank deltas, or streamin

Safe In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2509.25582v3 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks w

SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.28136v1 Announce Type: new Abstract: Dense semantic segmentation is essential for autonomous driving, yet many multi-modal datasets lack pixel-level annotations. The Zenseact Open Dataset (

SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning

Model ReleasesDGX agent

arXiv:2602.01990v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, but real-world deployment requires them to con

SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping

Model ReleasesDGX agent

arXiv:2605.28735v1 Announce Type: new Abstract: Transparent objects are common in daily life, and it is important to understand their multilayer depth, including the transparent surface and the object

Self-Supervised Online Robot-Agnostic Traversability Estimation for Open-World Environments

Model ReleasesDGX agent

arXiv:2605.28442v1 Announce Type: cross Abstract: Self-supervised online traversability estimation enables robots to continuously learn from unlabeled open-world experiences and adapt their navigation

SHIPPED. Mistral Vibe is now the AI agent for long-horizon productivity and coding, and the home for Work mode, Code mode, the CLI, and a br…

Model ReleasesDGX agent

Mistral AI has released Mistral Vibe, an AI agent designed for long-horizon productivity and coding tasks, featuring Work mode, Code mode, a CLI, and additional capabilities. The product consolidates

SIGMA: Bridging Structural and Distributional Gaps for Vision Foundation Model Adaptation

Model ReleasesDGX agent

arXiv:2605.27893v1 Announce Type: new Abstract: Vision Foundation Models (VFMs) have demonstrated impressive representational capabilities. However, adapting them to downstream tasks via full fine-tun

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

Model ReleasesDGX agent

arXiv:2605.28149v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) extract interpretable features from Large Language Models, but standard variants enforce non-negativity, forcing separate lat

Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering

Model ReleasesDGX agent

arXiv:2605.27636v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate excellent capabilities and performance for general reasoning tasks within the general public domain, t

SkillGrad: Optimizing Agent Skills Like Gradient Descent

Model ReleasesDGX agent

arXiv:2605.27760v1 Announce Type: new Abstract: Agent skills provide a lightweight way to adapt LLM agents to specialized domains by storing reusable procedural knowledge in structured files. However,

SmartIterator: Visual Analytics Workflows for Supervising Unsupervised Data Grouping

Model ReleasesDGX agent

arXiv:2605.28219v1 Announce Type: cross Abstract: Unsupervised learning methods -- topic modeling, partition-based and density-based clustering -- produce data groupings without human guidance, yet ch

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

Model ReleasesDGX agent

arXiv:2605.28122v1 Announce Type: cross Abstract: A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while th

Snippet-Driven Supply Chain Discovery with LLMs: Scaling Visibility in China

Model ReleasesDGX agent

arXiv:2605.27845v1 Announce Type: cross Abstract: Financial and economic research often relies on structured supply-chain disclosures and commercial databases. In China, supplier--customer disclosure

Snowveil: A Framework for Decentralised Preference Discovery

Model ReleasesDGX agent

arXiv:2512.18444v2 Announce Type: replace-cross Abstract: Aggregating subjective preferences in social choice traditionally assumes a trusted central authority. In contrast, this paper formalises Dece

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

Model ReleasesDGX agent

arXiv:2601.21666v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are a major focus of recent AI research. However, most prior work focuses on static image understanding, wh

Soro: A Lightweight Foundation Model and Chatbot for Tajik

Model ReleasesDGX agent

arXiv:2605.27379v1 Announce Type: new Abstract: We present Soro, a family of Tajik-specialized conversational large language models (LLMs) designed for real-world deployment under tight compute and co

Sparse POD Mode Selection and Manifold Dimensionality Reduction with Neural Networks

Model ReleasesDGX agent

arXiv:2605.27756v1 Announce Type: cross Abstract: High-performance computing enables simulation of high-dimensional physical systems, but downstream analyses such as inverse problems and control remai

SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs

Model ReleasesDGX agent

arXiv:2605.28490v1 Announce Type: cross Abstract: 3D object grounding localizes referred objects in a 3D scene from natural language. Unified instance-centric 3D-LLMs aim to solve grounding together w

Stay Fair! Ensuring Group Fairness in Diffusion Models Across Guidance Scales

Model ReleasesDGX agent

arXiv:2605.28036v1 Announce Type: new Abstract: Diffusion models steer conditional generation with a tunable guidance scale to trade off prompt alignment and diversity. However, existing debiasing tec

Stochastic Gradient Descent with Momentum is Algorithmically Stable

Model ReleasesDGX agent

arXiv:2605.28517v1 Announce Type: cross Abstract: Stochastic gradient descent with momentum (SGDM) is one of the most widely used optimization algorithms in machine learning. While optimization proper

StoryLens: Preference-Aligned Story Rewriting via Context-Aware Narrative Enrichment

Model ReleasesDGX agent

arXiv:2605.28073v1 Announce Type: cross Abstract: Story rewriting aims to adapt existing narratives to diverse reader preferences while preserving plot consistency and narrative coherence. Unlike conv

StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation

Model ReleasesDGX agent

arXiv:2605.27393v1 Announce Type: cross Abstract: Large language models (LLMs) can generate fluent dialogue, but prior works lack situational grounding, dynamic strategy control, and evaluation aligne

STR Robot: Design of an Autonomous Mobile Robot from Simulation to Reality

Model ReleasesDGX agent

arXiv:2605.28110v1 Announce Type: new Abstract: With the rapid development of simulation tools, the development and validation of autonomous robotic systems have become more efficient before real-worl

Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval

Model ReleasesDGX agent

arXiv:2605.11325v2 Announce Type: replace-cross Abstract: Every major benchmark for LLM memory systems, LoCoMo foremost, measures whether a model answered correctly, not whether the memory system retr

SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats

Model ReleasesDGX agent

arXiv:2605.27911v1 Announce Type: new Abstract: Suicide is a critical global public health challenge, causing approximately 720,000 deaths each year and calling for timely, effective prevention strate

🚨super bad news for three of the biggest IPOs in history:

Model ReleasesDGX agent

🚨super bad news for three of the biggest IPOs in history: Companies are starting to question whether soaring AI spending is delivering meaningful returns. An AI consultant tells us a client recently s

Super excited to finally share Dynamic Workflows in Claude Code!! We built this a couple months ago, and it has slowly become a daily driver…

Model ReleasesDGX agent

Super excited to finally share Dynamic Workflows in Claude Code!! We built this a couple months ago, and it has slowly become a daily driver for a bunch of people at Anthropic. A few tips for getting

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

Model ReleasesDGX agent

arXiv:2605.28179v1 Announce Type: new Abstract: Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream b

Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language

Model ReleasesDGX agent

arXiv:2605.27886v1 Announce Type: new Abstract: Tactile sensing is essential for robots to achieve human-like gentle manipulation. However, existing Vision-Language-Action (VLA) models struggle to exp

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

Model ReleasesDGX agent

arXiv:2605.27431v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across dive

TCP-MCP: Landscape-Guided Co-Evolution of Prompts and Communication Topologies for Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.27850v1 Announce Type: new Abstract: Effective multi-agent systems cannot be designed by selecting prompts or communication graphs in isolation. Agent behavior depends on the information an

Techno Week

Model ReleasesDGX agent

Techno Week is a thematic event or content series from Cohere, a leading AI company, likely featuring updates on technological innovations, product announcements, or industry insights shared via their

The Abstraction Gap in Vision-Language Causal Reasoning

Model ReleasesDGX agent

arXiv:2605.28779v1 Announce Type: new Abstract: Vision-language models (VLMs) generate fluent causal explanations, but current evaluations cannot distinguish linguistic plausibility from faithful caus

The Alignment Floor: When Persona Customization Is Safe

Model ReleasesDGX agent

arXiv:2605.27382v1 Announce Type: cross Abstract: A key promise of pluralistic AI is behavioral adaptation: persona prompts like 'be creative' or 'be thorough' let systems respect diverse user values

The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment

Model ReleasesDGX agent

arXiv:2605.28464v1 Announce Type: cross Abstract: Legal Judgment Prediction (LJP) has become a core benchmark for evaluating AI in the criminal legal domain, but it only sees criminal cases that have

The CFTC moves to vacate a $5M settlement with Gemini, reversing a Biden-era enforcement action, following a lobbying campaign by the Winklevoss twins (Wall Street Journal)

Model ReleasesDGX agent

Wall Street Journal: The CFTC moves to vacate a 5M settlement with Gemini, reversing a Biden-era enforcement action, following a lobbying campaign by the Winklevoss twins — A 5 million settlement at t

The European Commission launches a full review of JD.com's €2.2B acquisition of German electronics retailer Ceconomy under its Foreign Subsidies Regulation (Bloomberg)

Model ReleasesDGX agent

Bloomberg: The European Commission launches a full review of JD.com's €2.2B acquisition of German electronics retailer Ceconomy under its Foreign Subsidies Regulation — Chinese e-commerce firm JD.com

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

Model ReleasesDGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

Model ReleasesDGX agent

arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding rob

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

Model ReleasesDGX agent

arXiv:2605.28700v1 Announce Type: new Abstract: The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-

The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates

Model ReleasesDGX agent

arXiv:2605.28020v1 Announce Type: new Abstract: With the rapid progress of large language models (LLMs), reliably evaluating the capabilities of pre-trained LLMs has become increasingly important. The

The Name’s Gaming … Cloud Gaming: ‘007 First Light’ Launches on GeForce NOW

Model ReleasesDGX agent

License to stream, shaken and stirred. GeForce NOW is dialing up the espionage with the launch of 007 First Light, letting members slip into James Bond’s reimagined origin story from almost any device

The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study

Model ReleasesDGX agent

arXiv:2504.04540v2 Announce Type: replace-cross Abstract: 3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

Model ReleasesDGX agent

arXiv:2601.17737v3 Announce Type: replace-cross Abstract: Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, th

Thermodynamic properties of chemically disordered compounds via AI-driven estimation of partition function with the PULSE method

Model ReleasesDGX agent

arXiv:2605.28594v1 Announce Type: cross Abstract: In this article, we present an improved version of the PULSE method (Partition function Unsupervised Learning Sampling and Evaluation) for estimating

this is so funny, training opus 4.7 on business skills makes it misaligned and dishonest 😭

Model ReleasesDGX agent

this is so funny, training opus 4.7 on business skills makes it misaligned and dishonest 😭 Learnings from testing Claude Opus 4.8: > Much worse than Opus 4.7 and GPT 5.5 on Vending Bench > More aligne

Too many business leaders believe that AI says what it means. And it’s odd because we naturally attribute a high number of human traits to A…

Model ReleasesDGX agent

Too many business leaders believe that AI says what it means. And it’s odd because we naturally attribute a high number of human traits to AI, and yet we refuse to believe it can have hidden intent? 3

Took some inspiration from @vboykis and converted my first ever talk into a blog post. I talk about the role of agentic search in context en…

Model ReleasesDGX agent

Took some inspiration from @vboykis and converted my first ever talk into a blog post. I talk about the role of agentic search in context engineering. Together we build an intuition on the strengths a

Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution

Model ReleasesDGX agent

arXiv:2605.28000v1 Announce Type: cross Abstract: Large language model agents are increasingly expected to perform operational work: calling APIs, manipulating files, assembling workflows, and acting

Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

Model ReleasesDGX agent

arXiv:2605.27879v1 Announce Type: new Abstract: Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make ex

TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.28699v1 Announce Type: new Abstract: Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain d

Trinity: Unifying Class-Agnostic Terrain and Semantic Segmentation for Unstructured Outdoor Environments by Leveraging Synthetic Data

Model ReleasesDGX agent

arXiv:2605.27644v1 Announce Type: cross Abstract: Terrain understanding is fundamental for mobile robots operating in unstructured outdoor environments. Existing vision-based traversability estimation

Understanding Generalization and Forgetting in In-Context Continual Learning

Model ReleasesDGX agent

arXiv:2605.28705v1 Announce Type: new Abstract: In-context learning (ICL) derives its power from enabling Large Language Models to adapt to new tasks via prompt-based reasoning alone, entirely bypassi

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

Model ReleasesDGX agent

arXiv:2605.28063v1 Announce Type: cross Abstract: Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. C

UniMaia: Steering Chess Policies with Language for Human-like Play

Model ReleasesDGX agent

arXiv:2605.27767v1 Announce Type: cross Abstract: Recent advances in large language models have enabled natural language to serve as a flexible interface for controlling complex systems, but often at

Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis

Model ReleasesDGX agent

arXiv:2605.27401v1 Announce Type: cross Abstract: There is a growing interest in utilizing synthetic populations for a diverse range of applications. At the same time, we are witnessing a tremendous g

Verifiable Benchmarking of Long-Horizon Spatial Biology

Model ReleasesDGX agent

arXiv:2605.28065v1 Announce Type: new Abstract: AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or

← Previous
1…202203204205206…377
Next →