AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,098 results
11 Aug 2026

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Model ReleasesDGX agent

arXiv:2608.09548v1 Announce Type: cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that or

ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making

Model ReleasesDGX agent

arXiv:2608.09024v1 Announce Type: new Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prior

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2608.09189v1 Announce Type: new Abstract: Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Model

EndoMD-SLAM: Endoscopic Gaussian Splatting SLAM under Optical Degradation with Memory and Static-Transient Decomposition

Model ReleasesDGX agent

arXiv:2608.08949v1 Announce Type: new Abstract: Dense 3D reconstruction is critical for clinical endoscopic navigation and documentation. While Gaussian Splatting SLAM systems show promise in this dom

EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

Model ReleasesDGX agent

arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity

Entropy-based Code Adversarial Translation for Real-world Repository Migration

Model ReleasesDGX agent

arXiv:2608.09273v1 Announce Type: new Abstract: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnabl

Evaluating Generative Time-Series Models on Data with Point Masses

Model ReleasesDGX agent

arXiv:2608.09692v1 Announce Type: cross Abstract: Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no rid

Evo-Bench: Can Language Models Improve Agent Harness?

Model ReleasesDGX agent

arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emergi

EvoTrustRAG: Evolution-Aware Conflict Attribution and Evidence Handling for Reliable Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2608.07933v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) improves the factuality of large language models with external knowledge, yet conflicting evidence remains a fund

Exact Contraction Rates via the Berkson--Porta Representation: A Sharp Threshold and Its Herglotz-Kernel Obstruction

Model ReleasesDGX agent

arXiv:2608.07552v1 Announce Type: cross Abstract: Semigroups of holomorphic self-maps of the unit disc with an interior fixed point are, by the classical Berkson--Porta representation, entirely determ

Expert-Guided Multimodal Fusion for Unified Emotion and Sentiment Analysis

Model ReleasesDGX agent

arXiv:2601.07565v2 Announce Type: replace-cross Abstract: Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, while s

ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document type…

Model ReleasesDGX agent

ExtractBench is one of the most comprehensive benchmarks for real-world document extraction. ✅ It covers 4869 pages, across 67 document types, spanning 8 real-world domains: finance, energy, gov, auto

Failure-Mechanism Transferability of Cumulative-Damage Features for Health State Estimation of SiC Power Modules

Model ReleasesDGX agent

arXiv:2608.08365v1 Announce Type: cross Abstract: Data-driven health-state estimators for SiC (Silica-Carbide) power modules typically report their performance on a single accelerated-aging campaign,

Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders

Model ReleasesDGX agent

arXiv:2608.08284v1 Announce Type: new Abstract: Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable in

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Model ReleasesDGX agent

arXiv:2608.09474v1 Announce Type: new Abstract: Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained prima

FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models

Model ReleasesDGX agent

arXiv:2511.18852v2 Announce Type: replace Abstract: Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on gener

FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients

Model ReleasesDGX agent

arXiv:2608.09250v1 Announce Type: new Abstract: Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per

FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning

Model ReleasesDGX agent

arXiv:2608.09208v1 Announce Type: cross Abstract: Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learning (DFL). Ho

FemWear: A Specialized Wearable Foundation Model for Women's Health

Model ReleasesDGX agent

arXiv:2608.08244v1 Announce Type: new Abstract: General wearable foundation models are pretrained across broad sensor streams and populations, but are not designed around women's-health tasks. We intr

Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

Model ReleasesDGX agent

arXiv:2608.08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fi

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compa

FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) i

Floating-Point Neural Network Verification at the Software Level

Model ReleasesDGX agent

arXiv:2510.23389v2 Announce Type: replace-cross Abstract: The behaviour of neural network components must be proven correct before deployment in safety-critical systems. Unfortunately, existing neural

Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains

Model ReleasesDGX agent

arXiv:2608.07474v1 Announce Type: new Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cognitive

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

Model ReleasesDGX agent

arXiv:2505.16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of label

ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

Model ReleasesDGX agent

arXiv:2608.08605v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common ba

Forgetting-Resistant and Lesion-Aware Source-Free Domain Adaptive Fundus Image Analysis with Vision-Language Model

Model ReleasesDGX agent

arXiv:2602.19471v2 Announce Type: replace Abstract: Source-free domain adaptation (SFDA) aims to adapt a model trained in the source domain to perform well in the target domain, with only unlabeled ta

FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification

Model ReleasesDGX agent

arXiv:2608.08207v1 Announce Type: cross Abstract: Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales while rema

FriskAI launches with $3.6M to show enterprises what their AI agents are doing

Model ReleasesDGX agent

Runtime intelligence startup FriskAI Inc. launched today with 3.6 million in pre-seed funding to give enterprises a record of what their artificial intelligence agents actually do once they go into pr

From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.07770v1 Announce Type: cross Abstract: Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial con

From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing

Model ReleasesDGX agent

arXiv:2608.09842v1 Announce Type: new Abstract: Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on comp

From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations

Model ReleasesDGX agent

arXiv:2608.07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language

From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings

Model ReleasesDGX agent

arXiv:2608.08896v1 Announce Type: new Abstract: Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited access to special

From Product Search to Preference Articulation: The Economics of Agentic Commerce

Model ReleasesDGX agent

arXiv:2608.08395v1 Announce Type: cross Abstract: Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents. We compare

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Model ReleasesDGX agent

arXiv:2608.09925v1 Announce Type: cross Abstract: Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of p

Full-bandwidth transformer

Model ReleasesDGX agent

arXiv:2608.08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each

Fusion Training for Mathematical Generalization in Large Language Models

Model ReleasesDGX agent

arXiv:2608.09893v1 Announce Type: cross Abstract: Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Model ReleasesDGX agent

arXiv:2608.08722v1 Announce Type: cross Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely

GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis

Model ReleasesDGX agent

arXiv:2608.09921v1 Announce Type: new Abstract: Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains such as power s

GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

Model ReleasesDGX agent

arXiv:2608.09493v1 Announce Type: new Abstract: Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains chal

Goal-oriented Navigation Instruction Generation with Tour Video Priors

Model ReleasesDGX agent

arXiv:2608.08596v1 Announce Type: new Abstract: Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily t

Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification

Model ReleasesDGX agent

arXiv:2603.19329v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can generate plausible code but offer limited guarantees of correctness. Formally verifying that implementations

Google’s Gemini AI app passes 1 billion monthly active users

Model ReleasesDGX agent

Google LLC’s Gemini artificial intelligence app has passed 1 billion monthly active users, making it the 14th product in the company’s history to reach that mark. The company announced the milestone t

Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference

Model ReleasesDGX agent

arXiv:2608.09225v1 Announce Type: cross Abstract: The key-value (KV) cache is the primary throughput optimization in modern large language model (LLM) inference, enabling prefix reuse across requests.

Gradient Under Microscope: Benchmarking Resource Utilization of Memory-Efficient Gradient Computation Methods

Model ReleasesDGX agent

arXiv:2608.08961v1 Announce Type: new Abstract: AI training's rising resource intensity is straining electricity supplies and carbon budgets, motivating systematic study of memory-efficient training o

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

Model ReleasesDGX agent

arXiv:2608.07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environment

GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views

Model ReleasesDGX agent

arXiv:2608.09270v1 Announce Type: cross Abstract: Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment

Model ReleasesDGX agent

arXiv:2608.08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts

HeatCast: A Benchmark for Neighborhood-Scale LST Forecasting across 124 U.S. Cities

Model ReleasesDGX agent

arXiv:2608.07640v1 Announce Type: new Abstract: Land Surface Temperature (LST) is a widely used satellite-derived measure of urban surface heat, but there is no shared benchmark for forecasting it at

Hidden Language Consistency Phenomena in Reasoning LLMs

Model ReleasesDGX agent

arXiv:2608.08447v1 Announce Type: cross Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended langu

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

Model ReleasesDGX agent

arXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harne

High Fidelity Capture, Reconstruction, and Transfer of Human Demonstrations for Robot-Assisted Bathing

Model ReleasesDGX agent

arXiv:2608.09127v1 Announce Type: new Abstract: Despite the demand for robots in high-value clinical tasks like bathing, contemporary systems still lack the safety and reliability required for complex

High-Layer Attention Pruning with Rescaling

Model ReleasesDGX agent

arXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional

Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays

Model ReleasesDGX agent

arXiv:2608.07609v1 Announce Type: cross Abstract: High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens ty

HOPPER: Learnable Hop Extraction for Linearized Graph Sequence Models

Model ReleasesDGX agent

arXiv:2608.09031v1 Announce Type: new Abstract: Graph neural networks typically propagate information through repeated message-passing layers, coupling the distance over which information travels with

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

Model ReleasesDGX agent

arXiv:2608.07861v1 Announce Type: new Abstract: Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasse

How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare

Model ReleasesDGX agent

arXiv:2608.07511v1 Announce Type: cross Abstract: Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have b

How to Ask the AI: A User Perspective Survey for Large Language Model Prompting

Model ReleasesDGX agent

arXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by t

I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Model ReleasesDGX agent

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

Model ReleasesDGX agent

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon

← Previous
1…56789…369
Next →