每日文献雷达:2026-08-13

今日 slides:research-radar-2026-08-13

每日文献雷达:2026-08-13

今日自动检索并筛选出 3 篇候选论文,通过结构化深度阅读生成以下分析。


flowchart TB
    subgraph TODAY["今日入选 2026-08-13"]
    P1["AgenticDataBench: A Comprehensive Benchmark for Da..."]
    P2["DBA-Bench: A Production-Fidelity Benchmark for LLM..."]
    P3["LLMs for Knowledge Graph Construction and Reasonin..."]
    end
    TODAY --> BLOG["博客深度阅读"]
    BLOG --> SLIDES["Marp 幻灯片"]
    style TODAY fill:#f0f4ff,stroke:#4a90d9
    style BLOG fill:#d4edda,stroke:#155724
    style SLIDES fill:#e2d9f3,stroke:#6f42c1

今日入选

  • AgenticDataBench: A Comprehensive Benchmark for Data Agents(score: 0.5975)
  • DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents(score: 0.5175)
  • LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities(score: 0.4075)

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society.

  • 摘要:Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society. Automating this process is essential to reducing labor-intensive efforts for data scientists and enabling scalable data-driven applications. Recently, large language model (LLM)-based data agents have emerged as a promising solution to automate data science workflows. However, the field lacks comprehensive benchmarks to rigorously evaluate these agents across diverse scenarios with fine-grained granularity. To address this gap, we propose AgenticDataBench, a comprehensive benchmark featuring realistic tasks spanning diverse domains with fine-grained ground-truth labels. This enables evaluations to capture the diversity and complexity o…

  • 方法·三元组:当前环境未配置 LLM API Key,无法生成结构化三元组分析。GitHub Actions CI 中会使用百炼 API 进行深度阅读,采用公众号 storytelling 风格输出。本地可通过设置 LLM_API_KEYLLM_BASE_URLLLM_MODEL 环境变量启用。

  • 实验:需确认数据集、指标和 baseline。

  • 风险:离线或工具降级时摘要可能不足,不能替代人工精读。

  • 后续动作:深读方法和实验设计,建议人工使用 /readpaper 精读


DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

  • 作者:Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng
  • 入选原因:topic keywords matched; published this month
  • 来源信息:链接:https://arxiv.org/abs/2607.22165

LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison.

  • 摘要:LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a running database); observation-space scale and complexity (causal diagnosis across thousands of time series, business logs, and concurrent activity); solution-space openness (multiple remediations with different operational trade-offs); and scenario complexity and coverage (faults cascading across internal mechanisms and operational domains). We present DBA-Bench, a benchmark addressing these gaps through production fidelity, outcome-first evaluation, and controlled scenario reproducibility. It uses instrumented PostgreSQL environments with active wo…

  • 方法·三元组:当前环境未配置 LLM API Key,无法生成结构化三元组分析。GitHub Actions CI 中会使用百炼 API 进行深度阅读,采用公众号 storytelling 风格输出。本地可通过设置 LLM_API_KEYLLM_BASE_URLLLM_MODEL 环境变量启用。

  • 实验:需确认数据集、指标和 baseline。

  • 风险:离线或工具降级时摘要可能不足,不能替代人工精读。

  • 后续动作:保留为快读候选,后续按主题相关性跟进


LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities

This paper presents an exhaustive quantitative and qualitative evaluation of Large Language Models (LLMs) for Knowledge Graph (KG) construction and reasoning.

  • 摘要:This paper presents an exhaustive quantitative and qualitative evaluation of Large Language Models (LLMs) for Knowledge Graph (KG) construction and reasoning. We engage in experiments across eight diverse datasets, focusing on four representative tasks encompassing entity and relation extraction, event extraction, link prediction, and question-answering, thereby thoroughly exploring LLMs’ performance in the domain of construction and inference. Empirically, our findings suggest that LLMs, represented by GPT-4, are more suited as inference assistants rather than few-shot information extractors. Specifically, while GPT-4 exhibits good performance in tasks related to KG construction, it excels further in reasoning tasks, surpassing fine-tuned models in certain cases. Moreover, our investigati…

  • 方法·三元组:当前环境未配置 LLM API Key,无法生成结构化三元组分析。GitHub Actions CI 中会使用百炼 API 进行深度阅读,采用公众号 storytelling 风格输出。本地可通过设置 LLM_API_KEYLLM_BASE_URLLLM_MODEL 环境变量启用。

  • 实验:需确认数据集、指标和 baseline。

  • 风险:离线或工具降级时摘要可能不足,不能替代人工精读。

  • 后续动作:保留为快读候选,后续按主题相关性跟进

检索说明

  • 检索层:arXiv(deepxiv)+ Semantic Scholar + Google Scholar,每日查询轮换,保证论文多样性。
  • 阅读层:优先获取 arXiv HTML 全文,使用结构化三元组 + 公众号 storytelling 风格拆解论文逻辑。
  • 深度阅读方法论参考 /readpaper 技能。
  • 自动分析用于雷达筛选,重要论文仍需人工复核。

每日文献雷达:2026-08-13
http://zkkk123.cn/2026/08/13/research-radar/2026-08-13-daily-research-radar/
Author
Ke Zhang
Posted on
August 13, 2026
Licensed under