AI知识库 @ai521
316 subscribers
21.8K photos
42 videos
19 files
838 links
@ai521 专注分享最实用的AI内容

🤖 AI教程(新手到进阶)
🧠 AI知识科普(大模型 / 提示词 / 自动化)
📰 AI资讯更新(每日最新AI动态)
📚 AI实战技巧(写作 / 绘画 / 编程 / 赚钱)
🔧 最新AI工具推荐

每天更新AI干货
长期做一个真正有价值的AI频道
Download Telegram
AI工具发布

近期,一款基于人工智能的新工具正式推出,专注于帮助开发者快速理解代码结构。该工具利用AI技术,能够从任何GitHub仓库或本地代码库自动生成Mermaid格式的架构图。Mermaid是一种文本驱动的图表语言,常用于可视化系统架构和依赖关系。服务提供免费试用,新用户可获得1000积分。这一功能有望减少开发者手动分析和绘制图表的时间,提升代码理解和团队协作效率,尤其对处理大型项目或初学者具有实用价值。随着AI在软件开发中的应用日益广泛,此类工具正逐渐成为优化编程工作流的重要辅助。 #AI #科技 #编程 #开发工具 #代码架构 #Mermaid #GitHub #软件开发 #效率提升
研究显示AI医疗问答准确率仅76%,引发应用担忧

一项由宾夕法尼亚州立大学研究人员主导的新研究表明,ChatGPT等大型语言模型对日常健康相关问题的回答准确率接近76%,这引发了人们对其在真实世界中面对用户时的可靠性担忧。研究团队希望了解普通用户如何使用AI处理健康问题及其准确性,发现尤其是在神经学和皮肤科等专业领域,AI工具在训练有素的医生手中可能比在患者手中效果更好。相关成果将在2026年加拿大蒙特利尔的ACM FAccT会议上发表。 为评估准确性及潜在危害,研究人员在宾夕法尼亚州立大学举办了一场名为“Diagnose-a-thon”的AI竞赛。34名参与者(包括教职工和学生)提交了212个由患者或医生视角撰写的、涉及真实或虚构健康问题的提示,并使用ChatGPT-4o、Gemini-1.5 Pro等四种大语言模型生成回答。随后,9位执业医师根据6分制量表评估了这些AI回答的准确性及其潜在危害。研究结果指出,虽然人工智能有望变革医疗健康,但在神经学等专业领域的实际应用,可能更适合在专业医生的指导下进行。 #人工智能 #医疗健康 #AI准确性 #大语言模型 #数字医疗 #科技研究 #医疗应用 #研究发现
微星发布Claw 8 EX AI Plus新一代游戏掌机

微星在2026年台北国际电脑展(Computex)前夕正式推出了Claw 8 EX AI Plus游戏掌机。这款设备摒弃了前代产品采用的英特尔Lunar Lake移动芯片,转而搭载了专为手持设备优化的处理器。微星强调,Claw 8 EX AI Plus是全球首款配备最新发布的英特尔Arc G3 Extreme处理器的掌机,该处理器集成了Xe3 GPU核心阵列,专注于提升游戏运行的图形性能与流畅度。此举不仅体现了微星在便携式游戏硬件领域的创新,也可能推动整个行业向更高效能的移动游戏解决方案迈进。 #微星 #Claw8EXAIPlus #游戏掌机 #IntelArcG3 #处理器 #科技新闻 #Computex2026 #AI #游戏硬件
Claude Code发布自愈功能更新,破解开发者六大难题

Anthropic对其AI编程工具Claude Code进行了最大规模更新,首发引入自愈功能。此次升级旨在解决开发者在使用AI编程智能体时常见的六大痛点,包括终端显示闪烁、交互不确定性、错误信息模糊、历史记录压缩问题、MCP连接不稳定以及文件处理异常导致会话崩溃。具体改进包括全新终端渲染器减少视觉干扰、实时流式传输展示AI思考过程、优化错误提示更具可读性、压缩功能带进度显示、强化MCP协议连接韧性以及自动检测并绕过文件异常。这些更新使Claude Code更稳定可靠,降低了开发摩擦,提升了工具与本地生态的集成度,标志着AI编程工具从单纯智能向工业级可靠系统转型。 #AI #编程 #ClaudeCode #Anthropic #开发者工具 #软件工程 #自愈功能 #技术更新
开发者发现 GPT-5.4 模型在 Codex 环境中自称为 GPT

近日,有开发者在 Codex 平台上使用时发现,一个被标记为 GPT-5.4 的模型,在交互过程中却自我标识为 GPT-5。这一现象引发了关于 OpenAI 模型版本管理和公开命名体系的讨论。目前尚不清楚这是否属于内部测试版本的标识未更新,或是模型自身信息输出的一个小错误。该发现显示了大型语言模型在快速迭代过程中,其命名与实际版本号之间可能出现的微妙差异,也反映了开发者社区对前沿技术细节的密切关注。 #AI #GPT #人工智能 #模型开发 #技术观察 #OpenAI #大模型 #科技新闻
AWS推出时间序列基础模型Chronos

随着大语言模型的成功,基础模型范式正扩展至时间序列领域。亚马逊云科技(AWS)于2025年10月发布了其Chronos模型家族的最新版本——Chronos-2。该模型属于时间序列基础模型,旨在通过在大规模、多样化的时间序列数据上进行预训练,从而像处理文本的GPT模型一样,能够“开箱即用”地解决各类下游任务,如预测、异常检测和分类。 与传统为特定问题从头构建专用模型的工作流相比,TSFM(如Chronos-2)带来了根本性变革。用户只需输入历史序列数据并设定预测范围,即可通过单次推理调用获得预测结果,大幅降低了尝试成本和冷启动门槛。这意味着即使面对小数据集,模型也能凭借其预训练中获得的先验知识提供有价值的见解。文章以建筑电力需求预测为例进行了说明,指出该范式使得非专业人员也能快速进行分析,加速了业务决策周期。此举标志着时间序列分析正迈向更通用、高效的“基础模型”时代。 #时间序列 #基础模型 #Chronos2 #AWS #AI应用 #预测分析 #技术趋势
维基百科资深编辑威胁罢工抗议裁员风波

维基百科作为互联网上备受信赖的知识平台,由全球志愿者无偿维护。上周,支持其运营的非营利组织维基媒体基金会突然裁撤了一个小型但关键的工程师团队,此举立即引发社区震动。志愿者编辑和贡献者认为,裁员不仅切断了基金会与编辑社群之间的紧密纽带,更涉嫌打压工会权益,引发广泛质疑。经过多日激烈讨论,数百名活跃的资深维基百科编辑已公开表态支持罢工行动,以表达不满。然而,在一个以自愿贡献为主的平台上,罢工如何具体实施仍是未知数,凸显了非营利组织与社区之间平衡治理的挑战。 #维基百科 #编辑罢工 #裁员争议 #工会权益 #互联网信任 #科技新闻
AI初创公司免费清洁服务以训练家用机器人

AI训练初创企业Shift近日推出一项特殊服务:为用户提供免费的家庭清洁。其核心模式是通过录制清洁人员工作过程的视频,来训练未来能够执行家务的机器人。具体而言,该公司会派遣人员上门进行全面清洁,包括擦拭、吸尘、除尘、整理和清洗等工作,同时对全程进行拍摄记录。这些获取的宝贵数据将用于人工智能模型的训练,以期开发出能完成复杂家务劳动的智能机器人。此举旨在利用真实场景数据加速家庭服务机器人的研发进程。 #AI #机器人 #初创公司 #免费服务 #科技新闻 #家庭清洁 #数据训练
VDF AI – Multi-agent AI orchestration with dynamic model routing

Contact us Sörmlandsvägen, 192 54 Sollentuna Stockholm Sweden (+46) 70 563 81 85 info@
How to Stress-Test LLM Judges Fairly

A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test | alphaXiv []( []( []( []( []( []( We're hiring Paper 4 / 11 Hide Tools ⌘ / Open Tools A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test Camilo Chacón Sartori*José H. García Catalan Institute of Nanoscience and Nanotechnology (ICN2), CSIC and BIST, Campus UAB, Bellaterra, Barcelona, Spain @ @ Code and materials: *Corresponding author Abstract Retrieval-augmented generation (RAG) sys- tems are often compared by asking a large lan- guage model (LLM) judge which answer is better. For multi-hop RAG, this has become a measurement problem as much as a modeling problem: the same score can reflect retrieval
quality, answer length, lexical overlap, or a sta- tistical test that ignores clustered data. We ask what happens when these choices are made ex- plicit. We propose a minimum measurement stan- dard for LLM-as-a-judge comparisons in RAG. The standard fixes the top-100 candidate pool, evidence budget, answer cap, generator, and prompt; it also requires pre-registered hypothe- ses, cluster-aware inference, an exact clus- ter sign-flip check when feasible, and second- judge replication. Clustered benchmarks can overstate progress; the field should adopt this standard. We stress-test it with Genetic Algo- rithm Decoder for Multi-hop Evidence Compo- sition (GADMEC), an evolutionary evidence selector, on 400 multi-hop questions in com- puter science/machine learning (CS/ML) and Materials Science. The protocol changes the empirical story.A binomial test makes all four semantic-baseline comparisons look sig- nificant; cluster-aware inference leaves only one Bonferroni-significant result. BM25 beats pure semantic GADMEC under the same bud- get, while a lexical-semantic hybrid recovers in CS/ML and narrows the Materials Science gap. 1 Introduction RAG systems are increasingly compared with pair- wise LLM-as-a-judge protocols: two answers are shown to a judge model, and the preferred answer is counted as a protocol is attractive because it is simple and cheap to also hides important choices. A method can look bet- ter because it selected better evidence, because it induced longer answers, because it matched lexi- cal cues that dense retrieval missed, or because the statistical test treated clustered examples as inde- pendent. In multi-hop RAG, these mechanisms are easy to mix. This paper asks whether the headline of a RAG comparison survives when those choices are con- trolled. In our experiment, the answer is mixed. A binomial test would make all four pre-registered semantic-baseline comparisons look significant. Cluster-aware inference leaves only one compar- ison significant after Bonferroni correction, with two more significant only before correction. The empirical story therefore depends not only on the selector, but also on the measurement protocol used to evaluate it. We make two claims. First, pairwise LLM-as-a- judge evaluation for multi-hop RAG is more fragile than standard reporting suggests. Second, the field should adopt a minimum measurement standard for cluster-structured LLM-as-a-judge benchmarks. The standard fixes evidence and answer budgets. It separates confirmatory from exploratory analy- ses through pre-registration and a public deviation log, uses cluster-aware inference, adds an exact sign-flip check when the cluster count permits it, and replicates headline results with a second judge. These components are not individually new. The contribution is to make them operate together in one controlled RAG comparison and show how the conclusions change. We stress-test the standard with Genetic Al- gorithm Decoder for Multi-hop Evidence Com- position (GADMEC), a Biased Random-Key Ge- 1 arXiv:2605.27789v1 [ ] 27 May 2026 []( " " "mailto: @ ")[](mailto: @ "mailto: @ ")[]( " " netic Algorithm (BRKGA; Gonçalves and Resende, 2011) for evidence subset selection. GADMEC is the instrument, not the protagonist. Its evolu- tionary search is separated from the decoder that enforces budget and diversity constraints, so a random-fitness ablation can test whether the fitness function contributes beyond the decoder machin- ery. Contributions. 1.
A fixed-budget evaluation design for LLM- as-a-judge comparisons in RAG: same top- 100 candidate pool, 2000-token evidence bud- get, 300-token answer cap, and generator set- tings for all methods. 2. A pre-registered analysis protocol:four Bonferroni-corrected primary hypotheses, an addendum for length matching and ablations, and a deviation log for post-hoc analyses. 3. A cross-domain benchmark: 687 arXiv pa- pers (2024–2026), a 3-level taxonomy, 10 cross-subfield combinations per area, 20 ques- tions per combination, and 200 contrastive multi-hop questions per area. 4. A methodological demonstration: the head- line of a multi-hop RAG comparison can change when inference respects the cluster binomial test would have reported all four primary tests as signifi- cant; wild-cluster bootstrap leaves only one Bonferroni-significant result. The same proto- col exposes a lexical-vs-semantic axis: BM25 beats pure semantic GADMEC, while a hy- brid recovers in CS/ML. The paper unfolds as follows. Section 2 situates fixed-budget evaluation relative to LLM-as-a-judge evaluation, multi-hop question answering (QA), subset selection, and pre-registration. Sections 3 and 4 define GADMEC, the baselines, and the controlled evaluation protocol. Section 5 reports the pre-registered results, length controls, content- distance diagnostics, and ablations. Sections 6–8 interpret the mechanisms, conclude, and state the limitations. 2 Related Work LLM-as-a-Judge Evaluation for RAG. Reference-free LLM-as-a-judge evaluation is now common in RAG formalised the setup for retrieval-augmented generation (Es et al., 2024), and recent surveys catalogue its variants (Li et al., 2025). The closest concern for this paper is length bias (Zheng et al., 2023; Dubois et al., 2024). We control both input and output budgets, then ask what signal remains after length matching. Multi-hop QA , 2Wiki- MultiHop, MuSiQue, and GRADE target chained- evidence multi-hop reasoning (Lee et al., 2025). Our questions have a different shape:contrastive composition(“how do X and Y differ on as- pect Z”). This makes evidence selection budget- sensitive because a good answer must cover several sub-aspects at once. Combinatorial subset selection in retrieval. Maximal Marginal Relevance (MMR; Carbonell and Goldstein, 1998) and Determinantal Point Pro- cesses (DPPs; Kulesza and Taskar, 2012) are clas- sical instances. Genetic algorithms have been ap- plied to RAG primarily for adversarial attack and corpus poisoning (Cho et al., 2024).We use a BRKGA constructively: the candidate pool is the search space, and the selected subset is the evidence plan. Pre-registration in -registration is rare in natural language processing (NLP). We follow social-science practice with a primary registration, timestamped addenda, and a deviation log separat- ing confirmatory and exploratory analyses. Together, these strands motivate an evaluation instrument rather than only a new retriever. The selector is allowed to vary; the evidence budget, answer cap, judge protocol, and evidential status of each analysis are not. 3 Method We first define the selector, then the controls that make the comparisons interpretable. Figure 1 sum- marizes the full pipeline. The key rule is simple: every method sees the same candidates and uses the same generation settings. Only the evidence plan changes. 3.1 GADMEC fitness GADMEC’s fitness over a candidate evidence plan P combines five components: f(P) =α COV+β DIV+γ COST+δ COH+ε SUB 2
[]( "Gonçalves and Resende")[]( "2011")[]( "Es et al.")[]( "2024")[]( "Li et al.")[]( "2025")[]( "Zheng et al.")[]( "2023")[]( "Dubois et al.")[]( "2024")[]( "Lee et al.")[]( "2025")[]( "Carbonell")[]( "and Goldstein")[]( "1998")[]( "Kulesza and Taskar")[]( "2012")[]( "Cho et al.")[]( "2024")[]( "Figure 1") 1. Corpus construction Taxonomy 3 levels 2 areas×8 subareas×5 topics arXiv search per combo 700 candidate combos 10 approved combos / area 5 TOP + 5 NICHO 687 papers chunked + embedded 2. Question generation Contrastive prompt “How do X and Y differ in. . . ” Sample 3 papers / combo ×2 paper triplets GPT-5.4-mini T= 0.7 200 questions / area multi-hop, no method names 3. Evaluation Retrieval top-100 chunks by cosine 4 evidence plans GADMEC / Greedy / MMR / BM25 GPT-5.4-mini answers T= 0, seed=42 Opus 4.7 judge (API) randomized A/B positions Figure 1: GADMEC pipeline. The protocol first builds a taxonomy-filtered corpus, then generates contrastive multi-hop questions, and finally evaluates four selectors under the same candidate pool, evidence budget, generator, answer cap, and judge (Claude Opus 4.7, randomised answer-A/answer-B positions). The only intended source of variation is the evidence selector. We re-judge the six headline comparisons with a second strong proprietary judge from a different provider (DeepSeek V4 Pro, thinking mode) to assess inter-judge robustness (§5.5). Here COV rewards mean query similarity, and DIV rewards pairwise dissimilarity among selected is a normalised token penalty,COH is centroid–query similarity, and SUB measures cover- age of GPT-5.4-mini query sub-aspects embedded with all-MiniLM-L6-v2. We use α= 0.30, β= 0.15, γ= 0.00, δ= 0.15, ε= 0.40, with sub-coverage threshold 0.40. The BM25 lexical component ζ is 0 for pure se- mantic GADMEC and activated only in the hybrid analyses. 3.2 BRKGA decoder BRKGA (Gonçalves and Resende, 2011) repre- sents each plan as random keys in[0,1]n and leaves feasibility to a decoder. The decoder sorts chunks by key and accepts them greedily until the plan reaches the 2000-token budget. It also enforces minimum query similarity 0.15, redundancy thresh- old 0.80, and at least three k-means thematic clus- search uses population size 20, elite fraction 0.24, elite-inheritance probability 0.70, at most 50 generations, early stopping after 15 stag- nant generations, and seed 42. 3.3 Baselines We compare against greedy top-k budget-fill, Max- imal Marginal Relevance (MMR) with λ= 0.5 (Carbonell and Goldstein, 1998), and the BM25 lexical retriever (k 1= 1.5, b= 0.75). All meth- ods operate on the same top-100 cosine-pre-filtered candidate pool. Win-rate differences therefore re- flect how methods choose evidence from the same pool, not whether one method had access to better candidates. 3.4 Question generation and judging Question generation uses a contrastive multi-hop prompt with GPT-5.4-mini at T= area has 10 combinations,with 20 questions per use GPT-5.4-mini at T= 0, seed 42, evidence-only prompting, and max_completion_tokens=300. Claude Opus 4.7 judges randomised answer-A/answer-B pairs in a single pass using a fixed pairwise prompt asks for a global preference based on factual correctness, completeness, evidence sup- port, clarity, specificity of evidence-backed claims, multi-source synthesis, and coverage of the ques- tion are allowed but reserved for answers that are truly equivalent; ties are recorded and excluded from the win-rate denominator. 4 Experimental Setup