AI知识库 @ai521
323 subscribers
22K photos
42 videos
19 files
847 links
@ai521 专注分享最实用的AI内容

🤖 AI教程(新手到进阶)
🧠 AI知识科普(大模型 / 提示词 / 自动化)
📰 AI资讯更新(每日最新AI动态)
📚 AI实战技巧(写作 / 绘画 / 编程 / 赚钱)
🔧 最新AI工具推荐

每天更新AI干货
长期做一个真正有价值的AI频道
Download Telegram
亚马逊SageMaker AI与MLflow发布模型监控架构,应对数据漂移

在生产环境中部署的机器学习模型,其准确性和有效性会因消费者行为变化、新产品发布、传感器技术升级等不可控因素而迅速下降,导致数据漂移和模型漂移。为帮助组织主动监控模型性能,亚马逊云科技介绍了一种基于开源Evidently库和Amazon SageMaker AI with MLflow的模型监控架构。该方案适用于分类和回归等判别式模型,可计算数据漂移和模型漂移,并将结果整合到自定义仪表盘、通过Slack等渠道发送警报,甚至触发自动模型重训练流水线。架构支持批处理和实时推理两种场景,实时端点需启用数据捕获,并可使用AWS Lambda函数按计划或触发运行监控代码。通过这一开源可定制的监控方式,组织能在模型精度下降前及时干预,降低负面影响,同时控制成本并灵活集成到现有观测流水线中。 #亚马逊云科技 #SageMaker #MLflow #模型监控 #数据漂移 #机器学习 #AI #AWS #开源
How AWS Finance teams reclaimed hundreds of hours with Amazon Quick

Every finance professional knows the drill. Monday morning arrives, and your Financial Planning and Analysis (FP&A) team disappears into data compilation. They pull numbers from multiple systems, reconcile sources, build charts, and write commentary. All to answer a question that should be straightforward: what happened with revenue last week, and why? Across AWS Finance, teams were spending hundreds of hours a month on exactly this kind of work. Not analysis. Not strategy. Getting the data ready so the real work could begin. Amazon Quick is a generative AI assistant that connects across all your enterprise data and applications, so business users can search, analyze, and take action through natural language. It handles the complexity of querying millions of rows, running advanced analytics, and automating re
curring workflows so your team doesn’t need to. In this post, we show how AWS Finance used chat agents and Flows in Quick to transform two of their most time-consuming workflows. Setting financial targets for strategic customers requires reconciling bottom-up forecasts from business teams with top-down projections from leadership. It also demands enough depth to catch the risks hiding beneath historical data. The team built an Amazon Quick chat agent that connects directly to enterprise data sources and delivers sophisticated insights through natural language conversation. The agent queries millions of rows across Amazon Redshift data tables instantly while also searching external data signals. Screenshot showing Quick presenting a scenario analysis and creating a 5-sheet Microsoft Excel worksheet. Here’s what changed: Before: Analysts could deep-dive roughly a third of strategic customers in the time available between bottoms-up inputs and when top-level targets are due. The rest got surface-level coverage. A single customer analysis consumed up to 6 hours of manual work, including extracting data, running models, and documenting findings. After: The Quick agent evaluates statistical forecasts, runs regression analysis, Monte Carlo simulations, and performs scenario modeling across multiple factors in approximately 10 minutes per customer. It surfaces risks and opportunities that manual analysis missed. The team now covers their entire customer portfolio with even greater depth than before. “We have expanded from deep-diving a third of our strategic customers to covering our entire portfolio. Our finance team now spends time on what matters: partnering with the business to drive revenue, not compiling data or writing complex queries.” What makes this work: An analyst asks a question in natural language: “Run an opportunity and risk assessment for our top strategic accounts.” Quick then queries millions of rows, runs advanced analytics, and synthesizes structured data with unstructured insights from field reports and pipeline data. The agent does bull versus bear analysis by reviewing accounts with upside potential based on contract renewal timing and pipeline strength, and flags accounts with risk exposure. These are insights that traditional models missed. Because there’s no coding barrier, every finance professional on the team becomes a data analyst. Teams customize agents for different regions or business units, and the insights refresh automatically. If target setting is a periodic deep dive, regular business reviews are the recurring ritual that occupies FP&A teams everywhere. At AWS, every week, insights on revenue performance need to be compiled, analyzed, and packaged for leadership. And every week, that preparation consumes an entire Monday. The same AWS Finance team solved this by deploying Amazon Quick chat agents specific to each geographic region, connected through Flows to automate workflows that run on a set cadence without manual intervention. Video showing a blank Revenue Performance Analysis Flow that helps automate weekly business review workflows. Here’s what changed: Before: Every Monday, FP&A analysts spent a full morning compiling data from multiple systems, analyzing trends, manually reaching out to sales leads for customer anecdotes, and preparing talk tracks so leaders could understand what happened with revenue and why. The process was manual, repetitive, and left little time for strategic work. After:
Quick runs the Flow automatically each Monday morning. Region-specific chat agents analyze revenue performance across multiple dimensions: by charge type, by customer segment, and by growth contribution. They prepare comprehensive insights with ready-to-use talk tracks for leadership. Fresh analysis is waiting before the workday begins. Quick doesn’t only report numbers. It connects structured data from financial systems with unstructured insights from field reports to get to the why behind the trends. It examines customers across over a dozen dimensions, identifies patterns, and flags anomalies with context. “These insights are prepared automatically every Monday morning. Our team now spends time on strategic priorities instead of compiling disparate data. We spend more time on the why and on driving business outcomes.” These two use cases share a common thread. In both, the bottleneck wasn’t analytical skill, it was data compilation. Data was scattered across systems. Getting a complete picture required hours of manual extraction before any real analysis could begin. Amazon Quick removes that bottleneck by connecting directly to enterprise data sources and letting finance professionals interact with their data through natural language. The result isn’t incremental efficiency. It changes how finance teams spend their time: Across these use cases, the AWS Sales and Marketing Finance team reduced target-setting time from 6 hours to approximately 10 minutes per customer deep dive. They also removed the manual Monday routine for weekly business review preparation entirely. The time reclaimed went directly back into strategic work: risk analysis, customer anecdote synthesis, and identifying opportunities for growth. You don’t need to face Amazon-scale complexity to benefit. Every finance team deals with fragmented data, recurring reporting cycles, and the tension between compiling numbers and actually using them. Amazon Quick is designed for business users. Finance professionals set up chat agents and automated workflows themselves, without engineering support. They customize agents for their specific needs, refine them through iteration, and expand them across the organization as results prove out. If your team is spending more time preparing insights than delivering them, that’s the gap Quick is built to close. Learn more about Amazon Quick for Finance . In the next post in this series, we will explore how AWS Finance teams are using Quick to automate cost optimization and streamline approval workflows, turning hours of manual analysis into minutes. Sindhu is a Senior Tech Product Marketing Manager at AWS, leading go-to-market strategy for Amazon Quick. With 15+ years across Amazon, Uber, and Google, she’s passionate about making tech marketing relatable, inclusive, and grounded in real customer value. Outside work, she enjoys playing with her dog, and brewing coffee from different origins. Sarah is a Principal Product Manager, leading AI tooling strategy for AWS Finance. With 13+ years at Amazon spanning operations, e-commerce, and machine learning, she’s passionate about building AI solutions that are practical, and grounded in real business impact. She’s the proud mom of two energetic children and one grumpy Italian Greyhound.
Telegram必备的搜索引擎,极搜JISOU帮你精准找到,想要的群组、频道、视频、音乐

👉 t.me/jisou2?start=a_8247614025
Cardr:基于Apple Intelligence的本地名片扫描应用 无需云端

由开发者Konstantin Mishukov推出的Cardr是一款完全在iPhone上运行的离线名片扫描软件。用户只需将摄像头对准名片,即可一键提取姓名、职位、公司、电话、邮箱、网址和地址等信息,并直接保存至通讯录。所有识别过程依托Apple Intelligence在设备端完成,无需注册账号、无需订阅、无需上传云端,名片数据始终留在手机本地。隐私方面,开发者声明不收集任何用户数据。Cardr支持iPhone 15 Pro及更新机型,需要iOS 26.0及以上系统。前3次扫描免费,之后可通过一次性支付8.99美元解锁无限次扫描。该应用为注重隐私的用户提供了一种安全便捷的数字化名片管理方案。 #Cardr #名片扫描 #AppleIntelligence #隐私保护 #iOS #本地处理 #设备端AI #科技新闻
A Production RAG Pipeline for PDFs: Relational Parsing, TOC Retrieval, Typed Answers

This article opens Part III of Enterprise Document Intelligence , a series that builds an enterprise RAG system from four bricks: document parsing, question parsing, retrieval, and generation. It is the first of two parts on the upgraded pipeline: this part upgrades each brick, one contract at a time, on the same paper and the same question as Article 1 (minimal RAG) . The second part, Composing the four RAG bricks into one pipeline, tested on real documents (link to come), wires them into one call and runs it on several real documents. 📓 Runnable companion notebooks are on GitHub : doc-intel/notebooks-vol1 . A hundred lines of Python wire four functions together: parse the PDF, parse the question, retrieve a few pages, ask a model. That pipeline returns the right answer on a clean question again
st a paper with a built-in table of contents. On a real corpus it breaks the first time the user types “positonal encodig” with two typos, the first time the document is a 200-page contract with no PDF outline, the first time the question asks for every exclusion instead of one, the first time downstream code wants a typed object instead of a string. Four bricks need an upgrade each before the pipeline ships. The running paper is a public 15-page arXiv submission, Attention Is All You Need . The question, “What are the options for positional encoding?” , comes in with two typos so the question-parsing brick has something to do. The output is what an enterprise user needs: a typed answer, with verbatim quotes tied to line ranges, and a full audit trail from question to citation. Pick a clean question on a clean PDF, run it through the simplest RAG pipeline: parse the PDF, extract keywords from the question, retrieve a few pages, ask an LLM. On the Attention paper with the question “What are the options for positional encoding?” , that pipeline returns “sinusoidal positional encoding and learned positional embeddings” with one contiguous span of pages cited. It worked, on a clean question, on a clean paper. Drop the same pipeline into an enterprise setting and four weak points appear quickly, one per brick: The four bricks below address those four weak points. Same paper, same question, real LLM calls. The output is a typed list with verbatim quotes, line ranges, and the full chain of decisions that produced it. The shape that survives the upgrade is the same four bricks Article 1 (minimal RAG) introduced. What changes is the contract per brick: what each one consumes, what it produces, and how the next one consumes it. The diagram below is the full contract: every brick, the outputs it produces, and which downstream brick consumes each one (including the parsing_summary side channel into both question parsing and generation). The per-brick subsections then zoom into each box. Each of the four bricks is upgraded in its own articles, worth reading for the full contract: Parsing runs once and turns the PDF into the small set of tables every later brick reuses. In: pdf_path , the PDF on disk. Out: line_df , page_df , toc_df , parsing_summary . What comes out, one row per unit: Article 1 (minimal RAG) parsed the PDF into one DataFrame called line_df , one row per visible line of text. Enough for keyword retrieval, not enough for anything else. Article 5 (document parsing) reframes parsing as building a small relational set : line_df stays, page_df aggregates lines to pages with their text, and toc_df carries the document’s native table of contents. The bootstrap chunk at the top of the article already produced the three DataFrames on the Attention paper. A look at each: The page_num and line_num on each row are what turn an answer into a citation. Drawn back onto the page they came from, the rows are literal: every recognized line is one box, and its line_num sits in the gutter. Aggregating the lines page by page gives page_df , the natural page unit that carries whole-page text and page-level context: toc_df carries the native outline; on the Attention paper it has three levels and twenty-two entries, one row per section with its title, level, and start page: Three tables, same shape contract, same numeric primary keys. Downstream bricks read what they need without re-parsing the PDF; Retrieval (Section 2.3) scans keyword hits and reads
toc_df to anchor on the right section, then sizes the context around it (the whole section, or a line window) at the granularity the question implies. page_df is the page-level scan unit, toc_df the map, line_df the atomic lines the anchor and window are cut from. parse_pdf actually returns more in the same dict (image regions, internal references, named objects) plus a parsing_summary carrying document-level metadata (doc type, language, page count, layout, typical fields, a short summary); this section focuses on the three tables retrieval reads, and parsing_summary returns in the second part as the side channel that travels into the LLM bricks. Question parsing turns the raw user string into a typed brief the next two bricks can act on. In: the raw question , the expert’s concept_keywords_df , and parsing_summary for document context. Out: a ParsedQuestion carrying intent , keywords , a RetrievalQuery brief, and a GenerationBrief . Article 1 (minimal RAG) called get_keywords_from_question on the clean question and got a corrected_question plus a short list of keywords back, in one LLM call. That single call does two jobs at once: it fixes the surface typos and pulls the content keywords. The corrected keywords come straight out of the parse, with no separate spell-checking pass bolted on afterward. If a keyword genuinely does not exist in the document, the recovery is retrieval’s feedback loop (Article 13, the workflow pipeline), which re-searches with the terms the document actually uses, not a blind snap onto the nearest corpus token. What this article adds on top of the parse is the expert vocabulary layer. The corrected keywords are expanded with the domain terms a practitioner would also search for, pulled from the expert’s concept_keywords_df . Asked about positional encoding , the expansion adds the two concrete methods, sinusoidal and learned , so retrieval matches them even though the question never named them. Article 6 (question parsing) bundles both layers into the full parse_question brick. The output is a typed ParsedQuestion Pydantic carrying the user’s intent , the extracted keywords , a RetrievalQuery sub-brief that retrieval consumes ( main_query , rewrites , anchor_keywords , section_hint , layout_hint ), and structural_hints for page / sheet / slide pinning when the question carries them. The question’s intent also fixes how much to keep around the anchor: the whole section when it maps to one, or a line window for a pinpoint fact, so retrieval knows the granularity, not just where to look. Article 7A (retrieval as filtering) develops this two-level anchor / context model in full. Section 1.2 of Article 6 (question parsing) develops the two derived briefs pattern: the same ParsedQuestion yields one brief for retrieval and a GenerationBrief for generation, each carrying only the fields its consumer can act on. Each step is explicit, so retrieval downstream can use all of it and an audit log can be reconstructed. The noisy question is the same one a frustrated user would type: The one call already corrected the typos: optoins and posiitional came back as clean, content keywords, with no second spell-checking pass bolted on. Now expand them with the expert’s vocabulary, the concept_keywords_df table that maps each topic to the terms a practitioner would also search for: { "original_question": "What are the optoins for posiitional encoding?", "keywords": ["positional encoding"], "expanded_keywords": ["positional
encoding", "sinusoidal", "learned"] } Two things happened in one pass. The keywords came back corrected: posiitional and optoins were typos, and the single parse_question call fixed them while pulling the content noun phrase, dropping framing words like options . If a keyword still did not exist in the document, retrieval’s feedback loop (Article 13, the workflow pipeline) would recover it, not a blind snap onto the nearest corpus token. The expanded keywords come from concept_keywords_df . Asked about positional encoding , the expansion adds the two concrete methods the paper uses, sinusoidal and learned , so retrieval matches them even though the question never named them. The table is small on purpose: each entry narrows on the topic, not a generic word like position that would pollute the search. Article 6 (question parsing) develops how it is built and maintained. For the question-parsing code in detail, see the three articles that develop the brick: Retrieval narrows the document down to the lines generation will read. It filters on the structured tables, it does not search a vector index. In: the RetrievalQuery brief (from question parsing), plus line_df and toc_df (from document parsing). Out: a RetrievalResult : the kept pages and filtered_line_df , the section or line window generation reads. How it works, in two phases: Article 1 (minimal RAG) ran one retrieval method, keyword matching on page_df , and kept the top three pages by match count. Article 7 (retrieval) reframes retrieval as a filter on the small relational set built in Section 2.1: narrow the candidate scope using structured tables before scoring keywords. The keyword method still runs, but toc_df opens a second signal. The natural way to use it is not substring matching on titles. The author of the document already grouped lines into sections and wrote a title for each. A small LLM call can read the whole TOC, reason about which section answers the question, and return its picks with a one-sentence rationale. Retrieval works in two phases (Article 7A, retrieval as filtering). First it finds the anchor : keyword hits, counted per TOC section, tell the router which sections carry the question’s terms; reason_on_toc reads the TOC plus those counts and picks the section. Then it sizes the context around that anchor, following the granularity the question implies (Section 2.2): the whole section for a listing or section question, or a line window around the match for a pinpoint fact. The section is the natural context unit; page_df is the coarse scan surface, line_df the atomic lines the window is cut from. Here, to stay incremental on Article 1 (minimal RAG), this article runs the two detectors and merges their pages; the production hybrid routes them into sections instead (Article 7 (retrieval), the section-first arbiter above). The keyword method runs first (cheap, no LLM, deterministic). The LLM TOC router runs next on the same toc_df . The union of their pages goes to generation. Article 7B (anchor detection) develops the LLM TOC router in detail (the prompt, the Pydantic output, why it beats substring matching on real documents). Article 7C (the LLM arbiter) completes the picture, ranking every candidate page from every detector in a single call. We use the standalone reason_on_toc here because it carries its weight on its own: the upgrade from substring is the single most impactful change a team running Article 1 (minimal RAG) ’s pipeline can make to
retrieval. Article 7 (retrieval) introduces the unified retrieve_context(question, line_df, *, method, top_k, ...) → RetrievalResult dispatcher that routes to keyword / embedding / TOC / hybrid behind a single signature, and returns a typed RetrievalResult instead of two tuples. We use the raw retrieve_pages and reason_on_toc here to keep this article incremental on top of Article 1 (minimal RAG); production code calls retrieve_context or the shared dispatch_page_retrieval helper. A useful check before trusting that table. The first matching line column scans line by line, one line at a time. That works for single-token keywords. A multi-word keyword like positional encoding can be broken across a line break in the PDF, with positional at the end of one line and encoding at the start of the next. Each line on its own contains neither phrase. The line-by-line scan finds nothing, even when the keyword is right there on the page. Count the misses across the document: The line is a PDF rendering artifact. The text the parser sees as one line is whatever fits between two visual line-break decisions made by the PDF generator. There is no semantic reason the unit of detection should map to that arbitrary boundary. The fix is to detect on a passage instead, where a passage is a small window of adjacent lines joined with a space. Any keyword that exists in the passage will be found, no matter where the line breaks fall. The page-level match_count was already correct, because retrieve_pages joins all lines on a page before scanning. The fix targets the line-grained helpers that the snippet column, the highlighting, and any later chunking step need. From here on, every helper that scans below the page level uses a passage window, not a single line. The LLM TOC router fills the other gap exposed above. Same toc_df from parsing, but the LLM reads it whole and reasons about which section answers the question, returning the picked section ids plus a one-sentence rationale. The function itself is short (format each TOC row, drop it in a prompt with the question, parse a typed SectionSelection back); Article 7B (anchor detection) shows it in full. Run on the noisy question against the paper’s 22 TOC entries: One LLM call, the whole TOC inside the prompt, a typed list of section_ids back with a sentence of reasoning. The substring matcher this article used to ship would have caught 3.5 Positional Encoding on this specific question because encoding appears in the title. It would have missed every question phrased differently from the author. “What happens if we exit early?” against a contract whose section is titled “Termination” : substring zero, LLM one. The cost is a small LLM call (a few thousand tokens for a typical TOC, a few hundred milliseconds), and it is cached forever on identical inputs. This view is what an expert reads. Not a flat page list with cosine scores, but the document’s own outline, with the parsed-question keywords marked where they land. Article 7C (the LLM arbiter) consumes exactly this shape as a structured brief, one row per candidate, and ranks them with per-candidate roles + reasons. That arbiter is an upgrade on top of the LLM TOC router walked above, kept out of scope here to stay incremental on Article 1 (minimal RAG). For the retrieval code in detail, see the three articles that develop the brick: Generation fills a typed schema from the retrieved lines. It is controlled execution against a contract, not free-form prose.