AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
4 Jun 2026

With the new memory system, you can review and steer what ChatGPT remembers through a memory summary, with more visibility and control over …

Model ReleasesDGX agent

OpenAI introduced a new memory system for ChatGPT that allows users to review and control what the AI remembers across conversations through a memory summary feature. This update provides enhanced tra

Wordle 1,810 4/6 ⬛🟨⬛⬛⬛ ⬛⬛🟨⬛🟨 🟨🟩⬛🟨🟨 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This entry documents a Wordle game result where the player solved puzzle #1,810 in 4 attempts, using the color-coded feedback system (black for incorrect letters, yellow for correct letters in wrong p

xAI has released a blog on Partnering with Vapi for Voice

Model ReleasesDGX agent

xAI announced a partnership with Vapi to integrate voice capabilities into xAI's AI systems and services. The collaboration aims to enhance conversational AI by leveraging Vapi's voice technology plat


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

XSSR: Cross-Domain Self-Supervised Representative Selection for Efficient Annotation in Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2606.04301v1 Announce Type: new Abstract: Acquiring labeled medical image data is resource-intensive and a challenge further exacerbated in cross-domain scenarios where source and target dataset

'Your AI Text is not Mine': Redefining and Evaluating AI-generated Text Detection under Realistic Assumptions

Model ReleasesDGX agent

arXiv:2606.04906v1 Announce Type: cross Abstract: Although it is generally agreed that AI-generated text poses a broad societal risk, there is no common understanding in the AI-generated text detectio

3 Jun 2026

20x Faster Training Data Reads with Alluxio and Ray Data: A Cross-Region Benchmark

Model ReleasesDGX agent

This benchmark demonstrates how integrating Alluxio with Ray Data achieves 20x faster training data read speeds for cross-region machine learning workloads on Anyscale's platform. The study shows perf

95% token reduction. 30x faster execution. 90%+ task completion. Today at #MSBuild, we announced a major shift to move reasoning upstream: P…

Model ReleasesDGX agent

95% token reduction. 30x faster execution. 90%+ task completion. Today at #MSBuild, we announced a major shift to move reasoning upstream: Pinecone Nexus now integrates directly with @Microsoft OneLak

A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature

Model ReleasesDGX agent

arXiv:2606.03609v1 Announce Type: cross Abstract: Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move. But for navigation, what matte

A Benchmark for Semi-supervised Multi-modal Crowd Counting

Model ReleasesDGX agent

arXiv:2606.03646v1 Announce Type: new Abstract: This paper constructs the first benchmark on semi-supervised multi-modal crowd counting. To lay the foundation for this unexplored task, we first formul

A Fast Methane Detection Pipeline on Board Satellites Based on Mag1c-SAS and LinkNet

Model ReleasesDGX agent

arXiv:2606.03675v1 Announce Type: new Abstract: Methane is a potent greenhouse gas, and detecting leaks early via hyperspectral satellite imagery can help climate change mitigation efforts. Meanwhile,

A New Framework for Cybersecurity Refusals in AI Agents

Model ReleasesDGX agent

arXiv:2606.02644v1 Announce Type: cross Abstract: Agentic scaffolds have dramatically improved LLM performance on complex, long-horizon tasks, yielding both broad benefits and amplified risks in domai

A Single-Loop Bilevel Deep Learning Method for Optimal Control of Obstacle Problems

Model ReleasesDGX agent

arXiv:2601.04120v2 Announce Type: replace-cross Abstract: Optimal control of obstacle problems arises in a wide range of applications and is computationally challenging due to its nonsmoothness, nonli

A workflow audit is no longer the best way to figure out how to use AI in your job. Despite the advice from AI labs, I'm more convinced, bec…

Model ReleasesDGX agent

A workflow audit is no longer the best way to figure out how to use AI in your job. Despite the advice from AI labs, I'm more convinced, because of AI's reasoning capabilities and long context horizon

Acceptance-Test-Driven Evaluation Protocols for Business-Centric LLM Systems

Model ReleasesDGX agent

arXiv:2606.02755v1 Announce Type: cross Abstract: Large language model (LLM) applications are increasingly expected to satisfy deterministic institutional requirements while relying on probabilistic g

Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Model ReleasesDGX agent

arXiv:2602.12430v4 Announce Type: replace-cross Abstract: The transition from monolithic language models to modular, skill-equipped agents marks a defining shift in how large language models (LLMs) ar

Alibaba releases Qwen3.7-Plus, a multimodal proprietary model with a 1M-token context window, costing $2 per 1M tokens, 60% less than text-only Qwen3.7-Max (Carl Franzen/VentureBeat)

Model ReleasesDGX agent

Carl Franzen / VentureBeat: Alibaba releases Qwen3.7-Plus, a multimodal proprietary model with a 1M-token context window, costing $2 per 1M tokens, 60% less than text-only Qwen3.7-Max — However, like

AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task

Model ReleasesDGX agent

arXiv:2606.03967v1 Announce Type: cross Abstract: We describe AlignAtt4LLM, an IWSLT 2026 simultaneous speech translation system for English to German, Italian, and Chinese. The system is a synchronou

AmbientEye: A Dataset for Pupil Segmentation under Natural Ambient Infrared Illumination

Model ReleasesDGX agent

arXiv:2606.03774v1 Announce Type: new Abstract: Eye tracking is essential for smart glasses, as it provides insight into user attention for ambient intelligence applications. However, most existing ey

An Asymptotic Theory of Chain-of-Thought in In-Context Learning

Model ReleasesDGX agent

arXiv:2606.03217v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermed

Analytical Evaluation of DCA Convergence Properties for Minimizing Prediction Functions of Gaussian RBF Support Vector Regression

Model ReleasesDGX agent

arXiv:2606.03559v1 Announce Type: new Abstract: For nonconvex optimization problems whose objective is the prediction function of a trained Support Vector Regression (SVR) model with the Gaussian radi

Another banger open-source release. Miso One is an 8B text-to-speech model with real emotional range, so voiceovers carry warmth, hesitation…

Model ReleasesDGX agent

Another banger open-source release. Miso One is an 8B text-to-speech model with real emotional range, so voiceovers carry warmth, hesitation, and excitement instead of sounding flat. It's purpose-buil

Any2Poster: Any-Source Poster Generation Across Modalities and Domains

Model ReleasesDGX agent

arXiv:2606.02915v1 Announce Type: new Abstract: Visual posters are a compact medium for communicating dense information, yet progress on automatic poster generation remains difficult to measure becaus

AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following

Model ReleasesDGX agent

arXiv:2606.03116v1 Announce Type: cross Abstract: The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated eval

APIC: Amortized Physics-Informed Calibration using Neural Processes

Model ReleasesDGX agent

arXiv:2606.03355v1 Announce Type: new Abstract: Physics models are inherently imperfect due to misspecified or missing mechanisms, resulting in systematic discrepancies between model predictions and r

ArrowFlow: Hierarchical Machine Learning in the Space of Permutations

Model ReleasesDGX agent

arXiv:2604.04087v2 Announce Type: replace Abstract: We introduce ArrowFlow, a machine learning architecture that operates entirely in the space of permutations. Its computational units are ranking fil

As AI gets better, it reveals an empty promise

Model ReleasesDGX agent

This week we've got tandem hands-ons with Google's new Gemini AI agent - Spark - from my colleagues David Pierce and Jay Peters. Their takeaways are similar: It's so effective that it's scary. Spark k

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation

Model ReleasesDGX agent

arXiv:2606.03175v1 Announce Type: new Abstract: Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an underspecified natural-language d

Assessing and Mitigating Miscalibration in LLM-Based Social Science Measurement

Model ReleasesDGX agent

arXiv:2605.11954v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in social science as scalable measurement tools for converting unstructured text into variables t

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics

Model ReleasesDGX agent

arXiv:2507.21638v2 Announce Type: replace Abstract: The development of reinforcement learning (RL) algorithms has been largely driven by ambitious challenge tasks and benchmarks. Games have dominated

ATLAS: A Large-Scale Evaluation Benchmark for Adversarial LiDAR Perception

Model ReleasesDGX agent

arXiv:2606.02924v1 Announce Type: new Abstract: Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and pot

Auditable Climate Risk Intelligence from Fragmented ESG Data: Deterministic Orchestration and Imbalance-Aware Learning for Scope 1-3 Validation

Model ReleasesDGX agent

arXiv:2606.02604v1 Announce Type: cross Abstract: ESG and climate risk data remain fragmented across heterogeneous Scope 1, Scope 2, and Scope 3 reporting environments, while conventional validation p

AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification

Model ReleasesDGX agent

arXiv:2606.03031v1 Announce Type: new Abstract: Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone

Auditing Engagement Incentives in the Kidfluencer Ecosystem: A Multimodal Weak Supervision Approach

Model ReleasesDGX agent

arXiv:2606.03173v1 Announce Type: cross Abstract: The rise of `kidfluencers' on YouTube has raised ethical concerns about child digital labor and exploitation. While emerging legislation attempts to r

AURA: Action-Gated Memory for Robot Policies at Constant VRAM

Model ReleasesDGX agent

arXiv:2606.02775v1 Announce Type: new Abstract: The KV-cache is the right memory for datacenters but the wrong memory for robots. Datacenter inference batches many short requests and resets them, amor

Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging

Model ReleasesDGX agent

arXiv:2606.02809v1 Announce Type: new Abstract: Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation con

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes

Model ReleasesDGX agent

arXiv:2606.02724v1 Announce Type: cross Abstract: Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric

BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

Model ReleasesDGX agent

arXiv:2606.02798v1 Announce Type: new Abstract: Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited. Existing benchmarks

Benchmarking Visual State Tracking in Multimodal Video Understanding

Model ReleasesDGX agent

arXiv:2606.03920v1 Announce Type: new Abstract: Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacit

BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots

Model ReleasesDGX agent

arXiv:2509.14636v2 Announce Type: replace Abstract: Scale-consistent ego-motion estimation is fundamental for autonomous ground robots. Bird's-Eye-View (BEV) representation naturally addresses the sca

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs

Model ReleasesDGX agent

arXiv:2606.03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a

Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

Model ReleasesDGX agent

arXiv:2606.03318v1 Announce Type: new Abstract: Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

Model ReleasesDGX agent

arXiv:2606.03829v1 Announce Type: new Abstract: Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and a

Calibration Data Trade-offs Across Capability Dimensions: Why Multi-Source Mixing Matters for High-Sparsity LLM Pruning

Model ReleasesDGX agent

arXiv:2606.03328v1 Announce Type: cross Abstract: Post-training pruning compresses large language models to high sparsity using a small unlabelled calibration set, and recent work has concluded that t

Can Factual Opinions Be Edited (Manipulated) in Large Language Models?

Model ReleasesDGX agent

arXiv:2606.03096v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly integrated into various domains, making knowledge editing techniques crucial yet potentially hazardous. Cu

Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams

Model ReleasesDGX agent

arXiv:2603.19250v2 Announce Type: replace Abstract: Evaluating language models in streaming environments is critical, yet underexplored. Existing benchmarks either focus on single complex events or pr

CANMOT: Class-Aware Noise Modeling for Multi-Object Tracking in Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.03590v1 Announce Type: new Abstract: Kalman filter (KF)-based multi-object tracking (MOT) remains a strong baseline for autonomous driving due to its strong performance, computational effic

CAPER: Clause-Aligned Process Supervision for Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.03327v1 Announce Type: cross Abstract: Text-to-SQL systems are typically evaluated by query-level execution correctness, but this terminal signal provides little guidance about which interm

Causal Neural Probabilistic Circuits

Model ReleasesDGX agent

arXiv:2603.01372v2 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) enhance the interpretability of end-to-end neural networks by introducing a layer of concepts and predicting

Causal Preference Elicitation

Model ReleasesDGX agent

arXiv:2602.01483v2 Announce Type: replace-cross Abstract: We propose causal preference elicitation, a Bayesian framework for expert-in-the-loop causal discovery that actively queries local edge relati

Characterizing Detectability in 3DGS Poisoning: A Stage-wise Benchmark

Model ReleasesDGX agent

arXiv:2606.03499v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has rapidly emerged as a leading representation for real-time novel view synthesis, but recent work shows it is vulnerable

Chatbots Output Meaningful (but Problematic) Language

Model ReleasesDGX agent

arXiv:2606.02973v1 Announce Type: new Abstract: Are utterances by AI chatbots meaningful? Concretely, if a user asks, say, Anthropic's agent Claude, 'What is the capital of Spain?' and Claude answers,

ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

Model ReleasesDGX agent

arXiv:2606.02802v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model struct

ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models

Model ReleasesDGX agent

arXiv:2606.03157v1 Announce Type: new Abstract: Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making

COD10K-C: Benchmarking Robustness of Camouflaged Object Detection Under Natural Image Corruptions

Model ReleasesDGX agent

arXiv:2606.02603v1 Announce Type: new Abstract: Camouflaged object detection has improved substantially, but most standard benchmarks evaluate models only on clean images. This is not realistic becaus

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

Model ReleasesDGX agent

arXiv:2606.03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks ca

Coherent Swap Regret and Channel-Proof Learning

Model ReleasesDGX agent

arXiv:2606.02655v1 Announce Type: cross Abstract: External regret certifies stability only against replacing one's behavior by a fixed alternative. In a quantum game, this misses a natural physical mo

Combining Statistical Features and Deep Encodings for Rehearsal-Based Class-Incremental Time Series Classification

Model ReleasesDGX agent

arXiv:2606.03292v1 Announce Type: cross Abstract: Many systems used in real-world environments require adding new categories and incorporating new information without forgetting what was previously le

CoMPAS3D: A Dataset and Benchmark for Interactive Motion

Model ReleasesDGX agent

arXiv:2507.19684v2 Announce Type: replace-cross Abstract: Socially interactive humanoid robots must engage with humans through their bodies, adapting in real time to a partner's movement, intent, and

Compress then Merge: From Multiple LoRAs into One Low-Rank Adapter

Model ReleasesDGX agent

arXiv:2606.03723v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) enables parameter-efficient specialization of foundation models, but the proliferation of task-specific adapters fragments ca

Consistent Yet Wrong: Evidence Insensitivity in Spatial Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.02742v1 Announce Type: new Abstract: Spatial reasoning is fundamental to robotics, autonomy, and embodied AI, yet modern vision-language models (VLMs) remain unreliable on metric distance q

← Previous
1…172173174175176…377
Next →