AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,646 results
11 Aug 2026

AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions

Model ReleasesDGX agent

arXiv:2506.17455v3 Announce Type: replace Abstract: Robust visual recognition in underwater environments remains a significant challenge due to complex distortions such as turbidity, low illumination,

Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory

SafetyDGX agent

arXiv:2406.14373v3 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science

Backward Compatibility in Tree-Based Explanations and Enhanced CART Algorithm

ApplicationsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.08674v1 Announce Type: new Abstract: In the operation of machine learning models, model update is a fundamental process that requires careful consideration of its impact on downstream decis

Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms

Model ReleasesDGX agent

arXiv:2508.16481v3 Announce Type: replace Abstract: Ensuring the safe use of agentic systems requires a thorough understanding of the range of malicious behaviors these systems may exhibit. In this pa

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. H

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

Model ReleasesDGX agent

arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. Howev

Defining Decentralization: An Ontological Perspective

AgentsDGX agent

arXiv:2608.09748v1 Announce Type: cross Abstract: Decentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as security, d

Distribution-Free Conformal Prediction for Steel Fatigue Strength: Marginal Validity Is Not Enough

SafetyDGX agent

arXiv:2608.07589v1 Announce Type: cross Abstract: Predicting fatigue failure in steel components experimentally is costly because it requires testing across multiple compositions and processing condit

DualCert: A Solver for the Traveling Salesman Problem with Constraint-Coupled Learning

Model ReleasesDGX agent

arXiv:2608.09042v1 Announce Type: new Abstract: Large traveling salesman problem (TSP) instances require a solver to allocate limited computation while preserving the validity of its outputs. Existing

Efficient Cross-View Localization in 6G Space-Air-Ground Integrated Network

Local AiDGX agent

arXiv:2603.11398v2 Announce Type: replace-cross Abstract: Recently, visual localization has become an important supplement to improve localization reliability, and cross-view approaches can greatly en

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

Local AiDGX agent

arXiv:2608.08285v1 Announce Type: new Abstract: We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

ApplicationsDGX agent

arXiv:2608.07629v1 Announce Type: new Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu la

EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

Model ReleasesDGX agent

arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity

Epically Powerful: An open-source software and mechatronics infrastructure for wearable robotic systems

TutorialsDGX agent

arXiv:2511.05033v2 Announce Type: replace Abstract: Epically Powerful is an open-source robotics infrastructure that streamlines the underlying framework of wearable robotic systems - managing communi

EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac Sim

HardwareDGX agent

arXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the

Evo-Bench: Can Language Models Improve Agent Harness?

Model ReleasesDGX agent

arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emergi

FemWear: A Specialized Wearable Foundation Model for Women's Health

Model ReleasesDGX agent

arXiv:2608.08244v1 Announce Type: new Abstract: General wearable foundation models are pretrained across broad sensor streams and populations, but are not designed around women's-health tasks. We intr

flowengineR: A Modular and Extensible Framework for Fair and Reproducible Workflow Design in R

SafetyDGX agent

arXiv:2511.00079v2 Announce Type: replace Abstract: flowengineR is an R package designed to provide a modular and extensible framework for building reproducible algorithmic workflows for general-purpo

Forgetting-Resistant and Lesion-Aware Source-Free Domain Adaptive Fundus Image Analysis with Vision-Language Model

Model ReleasesDGX agent

arXiv:2602.19471v2 Announce Type: replace Abstract: Source-free domain adaptation (SFDA) aims to adapt a model trained in the source domain to perform well in the target domain, with only unlabeled ta

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

SafetyDGX agent

arXiv:2608.09158v1 Announce Type: cross Abstract: Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency

From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving

SafetyDGX agent

arXiv:2608.08941v1 Announce Type: cross Abstract: Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe wha

From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems

SafetyDGX agent

arXiv:2608.07627v1 Announce Type: new Abstract: Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documenta

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Model ReleasesDGX agent

arXiv:2608.09925v1 Announce Type: cross Abstract: Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of p

Fusion Training for Mathematical Generalization in Large Language Models

Model ReleasesDGX agent

arXiv:2608.09893v1 Announce Type: cross Abstract: Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Model ReleasesDGX agent

arXiv:2608.08722v1 Announce Type: cross Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely

High Fidelity Capture, Reconstruction, and Transfer of Human Demonstrations for Robot-Assisted Bathing

Model ReleasesDGX agent

arXiv:2608.09127v1 Announce Type: new Abstract: Despite the demand for robots in high-value clinical tasks like bathing, contemporary systems still lack the safety and reliability required for complex

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

Model ReleasesDGX agent

arXiv:2608.07861v1 Announce Type: new Abstract: Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasse

Introducing Unsloth Desktop app

Model ReleasesDGX agent

Hi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Su

Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025

Model ReleasesDGX agent

arXiv:2608.09280v1 Announce Type: new Abstract: Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promo

I’ve been thinking about how agents can learn inside world models for years. We decided to scale up our RSI Lab to bridge recursive self-imp…

TutorialsDGX agent

I’ve been thinking about how agents can learn inside world models for years. We decided to scale up our RSI Lab to bridge recursive self-improvement with physical AI and robotics. We are looking for f

JUMP-lite: Compact, reproducible benchmarking of cell representations

Model ReleasesDGX agent

arXiv:2608.07632v1 Announce Type: cross Abstract: Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Painting no

Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games

AgentsDGX agent

arXiv:2608.08268v1 Announce Type: cross Abstract: As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations r

Looker’s semantic layer governs Gemini Enterprise data for user trust

Model ReleasesDGX agent

For organizations deploying AI agents at scale, there’s often a critical divide between structured and unstructured data. While large language models (LLMs) excel at parsing text documents, emails, an

Machine Learning and Data Analysis Using Posets: A Survey

SafetyDGX agent

arXiv:2404.03082v3 Announce Type: replace Abstract: Partially ordered sets (posets) are discrete mathematical structures that formalize the notion of comparison without forcing every pair of objects t

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.08176v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-disti

MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks

Model ReleasesDGX agent

arXiv:2507.03162v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These mod

Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities

Model ReleasesDGX agent

arXiv:2608.08964v1 Announce Type: new Abstract: The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks

Model ReleasesDGX agent

arXiv:2507.19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a s

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

Model ReleasesDGX agent

arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information e

MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

Model ReleasesDGX agent

arXiv:2603.28590v3 Announce Type: replace Abstract: Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mis

Multi-Agent AI Safety as an Institutional Design Problem

SafetyDGX agent

arXiv:2608.09828v1 Announce Type: cross Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent wo

NTIRE 2026 Low-light Enhancement: Twilight Cowboy Challenge

ApplicationsDGX agent

arXiv:2608.09782v1 Announce Type: new Abstract: This paper presents a review of the NTIRE 2026 Low-light Enhancement: Twilight Cowboy Challenge. The objective of the competition was to merge a set of

Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction

Local AiDGX agent

arXiv:2608.08443v1 Announce Type: cross Abstract: Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other parts of a

QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing

AgentsDGX agent

arXiv:2608.07743v1 Announce Type: new Abstract: Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the ta

REFRAMED: Towards Realistic Audio Description Generation for Movies

Model ReleasesDGX agent

arXiv:2608.09765v1 Announce Type: new Abstract: Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video cap

Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access

Model ReleasesDGX agent

arXiv:2608.08942v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide

SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment

Model ReleasesDGX agent

arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill s

Source: Trajectory, founded by ex-DeepMind, Apple, OpenAI, and Meta staffers to build continual learning models, raised 40M led by Sequoia at a 300M valuation (Stephanie Palazzolo/The Information)

IndustryDGX agent

Stephanie Palazzolo / The Information: Source: Trajectory, founded by ex-DeepMind, Apple, OpenAI, and Meta staffers to build continual learning models, raised 40M led by Sequoia at a 300M valuation —

Stealing Reasoning Traces from Proprietary LLM APIs

SafetyDGX agent

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

SuperCoder: Assembly Program Superoptimization with Large Language Models

Model ReleasesDGX agent

arXiv:2505.11480v4 Announce Type: replace-cross Abstract: Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its inp

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

Model ReleasesDGX agent

arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

Model ReleasesDGX agent

arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution c

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

SafetyDGX agent

arXiv:2608.07980v1 Announce Type: cross Abstract: In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the

Three Generations of Healthcare IT: From the Digital Record to the Computable Care Process

ApplicationsDGX agent

arXiv:2608.08806v1 Announce Type: new Abstract: Objective. Healthcare IT is usually organized by the technologies it adopts. We instead organize it by the unit of information a system makes computable

ToolUniverse: An open platform for democratizing AI scientists

SafetyDGX agent

arXiv:2509.23426v3 Announce Type: replace Abstract: AI scientists are emerging computational systems that serve as collaborative partners in discovery. These systems remain difficult to build because

Towards an LLM-based method for quantifying the sexual content in song lyrics

Model ReleasesDGX agent

arXiv:2608.08885v1 Announce Type: cross Abstract: Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mo

Towards Expert-level Medical AI for Real-time Video Consultations

Model ReleasesDGX agent

arXiv:2608.09861v1 Announce Type: new Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through

Using the different models in different industries - your experience?

IndustryDGX agent

Hi all, I'm seeing so many conversations from people discussing how they're 'using the models wrong' and 'don't use Sol max as high is enough for you', yet all the conversations lack the nuance of wha

verdi: retrieval is not transfer for continual world model optimization

HardwareDGX agent

arXiv:2608.09537v1 Announce Type: new Abstract: Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained world model t

We built the Agentic World Cup - LLMs that compete in 1v1 Soccer. [P]

AgentsDGX agent

Hey everyone - we've been building something particularly relevant to ML at large - The Agentic World Cup - a platform where Agents compete in sports. As you know, today's Agents can code, do math, an

← Previous
1…372373374375376…428
Next →