diff --git a/apps/api/app/data/demo_documents/tsla-q4-2025/chunks.json b/apps/api/app/data/demo_documents/tsla-q4-2025/chunks.json index 177ae25c9..68d1d3111 100644 --- a/apps/api/app/data/demo_documents/tsla-q4-2025/chunks.json +++ b/apps/api/app/data/demo_documents/tsla-q4-2025/chunks.json @@ -25,7 +25,7 @@ "chunk_id": "68a6be7d-c587-5c73-abf2-56f4686e28e6", "type": "text", "content": "[tables/table-0 Tesla 2025 Results.html]", - "path": "TSLA-Q4-2025-Update.pdf-->HIGHLIGHTS", + "path": "TSLA-Q4-2025-Update.pdf/HIGHLIGHTS", "metadata": { "length": 83, "summary": "", @@ -138,7 +138,7 @@ "chunk_id": "60109008-6261-51e6-b202-093d904eb881", "type": "text", "content": "FINANCIAL SUMMARY\n(Unaudited)\n\n[tables/table-1 Q4 2025 Financials.html]\n\n(1) As a result of the adoption of the new crypto assets standard, the previously reported quarterly periods in 2024 have been recast.\n(2) Beginning in Q1'25, Adjusted EBITDA (non-GAAP) is presented net of digital assets gains and losses and all prior periods have been adjusted.\n(3) Beginning in Q1'25, Net income attributable to common stockholders (non-GAAP) is presented net of digital assets gains and losses and all prior periods have been adjusted.\n(4) Beginning in Q1'25, Capital expenditures is presented inclusive of purchases of energy generation and storage systems and all prior periods have been adjusted.\nFINANCIAL SUMMARY\n(Unaudited)\n\n[tables/table-2 Financial Data 2021-25.html]\n\n(1) Beginning in Q1'25, Adjusted EBITDA (non-GAAP) is presented net of digital assets gains and losses and all prior periods have been adjusted.\n(2) Beginning in Q1'25, Net income attributable to common stockholders (non-GAAP) is presented net of digital assets gains and losses and all prior periods have been adjusted.\n(3) Beginning in Q1'25, Capital expenditures is presented inclusive of purchases of energy generation and storage systems and all prior periods have been adjusted.\nOPERATIONAL SUMMARY\n(Unaudited)\n\n[tables/table-3 Tesla Q4-2025 Data.html]\n\nOPERATIONAL SUMMARY\n(Unaudited)\n\n[tables/table-4 Tesla 2021-2025 Data.html]", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY", "metadata": { "length": 1517, "summary": "The document presents unaudited financial and operational summaries for a company, covering quarterly data from Q4 2024 through Q4 2025 and annual data from 2021 to 2025. Key notes indicate significant accounting changes effective Q1 2025: Adjusted EBITDA and Net income attributable to common stockholders are now presented net of digital assets gains and losses, with all prior periods adjusted accordingly. Additionally, Capital expenditures now include purchases of energy generation and storage systems, requiring restatement of previous periods. The content references multiple tables detailing these metrics but does not display the specific numerical values.", @@ -240,7 +240,7 @@ "chunk_id": "60339310-5480-5ae2-8791-e6017bafb730", "type": "text", "content": "While automotive sales declined sequentially, gross margin (even when excluding the impact of regulatory credits) improved. The APAC region continued to show strength across multiple markets and set a record for deliveries in the quarter. We continued the rollout of Model Y variants across markets in Q4, including the standard and performance versions.\nPreparations continue in North America for the production ramps of Tesla Semi and Cybercab, both commencing 1H26, and production of the next-generation Roadster.", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->Automotive", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/Automotive", "metadata": { "length": 516, "summary": "", @@ -304,7 +304,7 @@ "chunk_id": "eb18e33e-b093-532e-b4df-bd05d63b0294", "type": "text", "content": "We achieved our highest quarterly energy storage deployments, driven by record Megapack deployments. Total gross profit rose, both sequentially and year-over-year, to a record \\$1.1 billion, marking the fifth consecutive record quarter. We plan to begin Megapack 3 and Megablock production at Megafactory Houston in 2026. In 2025, our global Powerwall network supported more than 89,000 Virtual Power Plant events across over 1 million installed units, allowing homeowners to save over \\$1 billion in electricity bills as Virtual Power Plant participation continues to scale rapidly.", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->Energy generation and storage", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/Energy generation and storage", "metadata": { "length": 583, "summary": "", @@ -395,7 +395,7 @@ "chunk_id": "60518074-bcbe-5b48-89a9-dfeb5d1b531b", "type": "text", "content": "We made further progress on the Optimus program in 2025. In Q1 of this year, we plan to unveil the Gen 3 version of Optimus, which will include major upgrades from version 2.5, including our latest hand design. The Gen 3 is our first design meant for mass production. Preparations are underway for the first production line, including supply chain readiness, with start of production planned before the end of 2026 and eventual planned capacity of 1 million robots per year.\nInstalled Annual Manufacturing Capacity\n\n[tables/table-5 Tesla Production.html]\n\nInstalled capacity ≠ current production rate and there may be limitations discovered as production rates approach capacity. Production rates depend on a variety of factors, including equipment uptime, component supply, downtime related to factory upgrades, regulatory considerations and other factors. Construction includes factory and infrastructure buildout as well as tool installation.", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->Robotics", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/Robotics", "metadata": { "length": 959, "summary": "", @@ -491,7 +491,7 @@ "chunk_id": "c49883a3-5834-5537-be53-35e58da70bf7", "type": "text", "content": "We are currently building Cortex 2 at Gigafactory Texas to further increase our AI training compute capacity. In the first half of 2026, we plan to more than double the size of onsite compute in Texas (in terms of H100 equivalents). We aim to maximize capital efficiency by scaling training compute judiciously, including when the training backlog gets too long or in anticipation of greater demand from our engineers to support our AI-related offerings.", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->AI Training Compute", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/AI Training Compute", "metadata": { "length": 454, "summary": "", @@ -545,7 +545,7 @@ "chunk_id": "fcbc3b02-dec6-5e91-849e-fccd3f22320e", "type": "text", "content": "Our lithium refinery commenced pilot production and is the first spodumene to lithium hydroxide refinery in North America, leveraging a simpler, cheaper and more environmentally friendly process. This refinery enables us to domestically produce critical minerals in support of energy storage, battery manufacturing and ultimately for EV growth.\nWe have begun to produce battery packs for certain Model Ys with our 4680 cells, unlocking an additional vector of supply to help navigate increasingly complex supply chain challenges caused by trade barriers and tariff risks. We now produce dry-electrode for 4680 cells with both anode and cathode made in Austin. We expect both domestic cathode material in Texas and LFP lines in Nevada to begin production in 2026.", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->Battery", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/Battery", "metadata": { "length": 762, "summary": "", @@ -667,7 +667,7 @@ "chunk_id": "34777c71-d716-5c39-b18a-bec0d0b64706", "type": "text", "content": "We continue to efficiently utilize our existing physical footprint in North America, with targeted augmentation to support the rollout of Robotaxi. While in the short-term, operational workstreams such as charging, cleaning and maintenance can be managed through our existing charging network, service centers and sales and delivery locations, we will have to add more capacity as the service expands. We added over 3,800 net new Supercharging stalls, growing the network 19% year-over-year.\nInstalled Annual Capacity\n\n[tables/table-6 Facility Status.html]\n\nInstalled capacity ≠ current production rate and there may be limitations discovered as production rates approach capacity. Production rates depend on a variety of factors, including equipment uptime, component supply, downtime related to factory upgrades, regulatory considerations and other factors. Construction includes factory and infrastructure buildout as well as tool installation.\n\n## Other Supporting Infrastructure Tesla AI Training Capacity Ramp (H100 equivalent GPUs)0\n[images/image-1-Capacity Growth Projection.jpg]\n\nTesla AI Training Capacity Ramp (H100 equivalent GPUs)", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->Other Supporting Infrastructure", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/Other Supporting Infrastructure", "metadata": { "length": 1142, "summary": "", @@ -786,7 +786,7 @@ "chunk_id": "778a63c2-7955-514c-84d0-9c2cfd99a489", "type": "text", "content": "We continue to enhance FSD (Supervised) $^{1}$ via our end-to-end foundation model trained on both customer and Robotaxi real-world data with our latest version, v14. FSD (Supervised) $^{1}$ increasingly provides safety and convenience functionality that can relieve drivers of many tedious and potentially dangerous aspects of road travel, including giving access to personal transport for those who otherwise have difficulty driving. V14 offers unparalleled driver assistance to safely drive to the customer's destination, find a free parking spot and park at that location. Our global fleet can collect the equivalent of over 500 years of continuous driving data per day $^{2}$ , allowing us to safely deploy and scale capabilities that can handle the long and fat tail of corner cases across diverse geographies and driving environments.", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->AI Software", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/AI Software", "metadata": { "length": 841, "summary": "", @@ -876,7 +876,7 @@ "chunk_id": "164babde-7a64-5b81-9a5a-41443b221c40", "type": "text", "content": "Development of our in-house, custom designed AI5 and AI6 inference chips for autonomy progressed during the quarter, with production planned for 2027 and 2028, respectively. We are targeting a 50x improvement in performance for AI5 relative to AI4 thanks to 10x raw compute, 9x memory capacity and 5x hardened block quantization and softmax function (the latter enabling efficient low-precision computing without sacrificing model accuracy).", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->AI Inference Compute", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/AI Inference Compute", "metadata": { "length": 441, "summary": "", @@ -970,7 +970,7 @@ "chunk_id": "32bfae6f-b2d4-51d3-94d9-82b9640828dd", "type": "text", "content": "The Robotaxi iOS app no longer has a waitlist in the areas we serve. Our vehicles keep getting better with our over-the-air updates, including: Grok (an AI companion) which now supports navigation commands (allowing users to find, add and edit navigation destinations hands-free); Tesla Photobooth which enables users to take photos in their car and download or share via the Tesla mobile app; Supercharger Site Maps which displays Supercharger layouts, nearby businesses and live availability details; Automatic HOV Lane Routing based on interior camera occupancy detection; Phone Left Behind Chime and SpaceX ISS Docking Simulator Game.\n\nWe continue to enhance FSD (Supervised) $^{1}$ via our end-to-end foundation model trained on both customer and Robotaxi real-world data with our latest version, v14. FSD (Supervised) $^{1}$ increasingly provides safety and convenience functionality that can relieve drivers of many tedious and potentially dangerous aspects of road travel, including giving access to personal transport for those who otherwise have difficulty driving. V14 offers unparalleled driver assistance to safely drive to the customer's destination, find a free parking spot and park at that location. Our global fleet can collect the equivalent of over 500 years of continuous driving data per day $^{2}$ , allowing us to safely deploy and scale capabilities that can handle the long and fat tail of corner cases across diverse geographies and driving environments. Cumulative Miles Driven with FSD (Supervised) $^{1}$ (billions)0\n[images/image-2-FSD Mileage Growth.jpg]\n\nCumulative Miles Driven with FSD (Supervised) $^{1}$ (billions)\n\nDevelopment of our in-house, custom designed AI5 and AI6 inference chips for autonomy progressed during the quarter, with production planned for 2027 and 2028, respectively. We are targeting a 50x improvement in performance for AI5 relative to AI4 thanks to 10x raw compute, 9x memory capacity and 5x hardened block quantization and softmax function (the latter enabling efficient low-precision computing without sacrificing model accuracy). Targeting Step-Function Improvement for our Next-Generation Inference Chip, AI50\n[images/image-3-Tesla Silicon Optimization.jpg]\n\nTargeting Step-Function Improvement for our Next-Generation Inference Chip, AI5\n(1) Active driver supervision required; does not make the vehicle autonomous\n(2) Calculated based on continuous hours of driving at an average of 30 miles per hour", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->Automotive and Other Software", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/Automotive and Other Software", "metadata": { "length": 2444, "summary": "The Robotaxi iOS app has removed its waitlist in served areas and now includes features like Grok navigation, Tesla Photobooth, Supercharger maps, automatic HOV routing, phone left-behind chimes, and a SpaceX game. FSD (Supervised) v14 uses an end-to-end foundation model trained on vast real-world data to assist drivers with navigation, parking, and safety, though active supervision remains required. The company is developing custom AI5 and AI6 inference chips for 2027 and 2028, targeting significant performance improvements over previous generations to handle complex driving scenarios globally.", @@ -1205,7 +1205,7 @@ "chunk_id": "41533301-58f0-554f-916d-8254cc2700df", "type": "text", "content": "We began testing driverless Robotaxis in Austin in December and began removing the safety monitor from customer rides in January on a limited basis, which will unlock further expansion of our Robotaxi fleet and coverage area in the Austin-metro. Our Bay Area ride-hailing service began serving the San Jose Airport in October, with plans to expand to other major airports in the Bay Area upon receiving required permitting.", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->Robotaxi", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/Robotaxi", "metadata": { "length": 423, "summary": "", @@ -1263,7 +1263,7 @@ "chunk_id": "3968a25f-933e-5a7f-a118-07db2e581153", "type": "text", "content": "We launched FSD (Supervised) $^{1}$ in South Korea, where customers drove over 1 million kilometers using the software in just one month. While we continue to pursue regulatory approval in China and Europe, we began offering ride-along experiences to consumers in Italy, Germany, France and Switzerland.\nMonthly subscriptions to FSD (Supervised) $^{1}$ continued to grow sequentially and more than doubled in 2025. Starting this quarter, we are transitioning access to FSD (Supervised) $^{1}$ to monthly subscriptions only as we begin to sunset the up-front payment option.", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->FSD (Supervised) $^{1}$", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/FSD (Supervised) $^{1}$", "metadata": { "length": 573, "summary": "", @@ -1364,7 +1364,7 @@ "chunk_id": "adf89909-d366-51e1-b68c-459f9024191c", "type": "text", "content": "Services and Other gross profit of approximately \\$300 million was partly driven by Part Sales and Supercharging. We now offer Tesla Insurance in Florida, as we continue to expand our insurance product to new states. In certain states, customers receive a discount on their insurance premiums when using FSD (Supervised) $^{1}$ . The more you drive with FSD (Supervised) $^{1}$ enabled, the bigger the discount is on your insurance premium – helping, in certain cases, to completely offset the monthly subscription cost for FSD (Supervised) $^{1}$ .\n\n## FSD (Supervised) $^{1}$ Cumulative Paid Robotaxi Miles0\n[images/image-4-Growth Trend 2025.jpg]\n\nCumulative Paid Robotaxi Miles\n\n[tables/table-7 Autonomous Driving Status.html]\n\nPlanned Robotaxi Coverage", - "path": "TSLA-Q4-2025-Update.pdf-->SUMMARY-->Automotive Services", + "path": "TSLA-Q4-2025-Update.pdf/SUMMARY/Automotive Services", "metadata": { "length": 742, "summary": "", @@ -1450,7 +1450,7 @@ "chunk_id": "cf91377e-176a-5768-855a-b859092f4695", "type": "text", "content": "On January 16, 2026, Tesla entered into an agreement to invest approximately \\$2 billion to acquire shares of Series E Preferred Stock of xAI as part of their recent publicly-disclosed financing round. Tesla’s investment was made on market terms consistent with those previously agreed to by other investors in the financing round. As set forth in Master Plan Part IV, Tesla is building products and services that bring AI into the physical world. Meanwhile, xAI is developing leading digital AI products and services, such as its large language model (Grok).\nIn that context, and as part of Tesla's broader strategy under Master Plan Part IV, Tesla and xAI also entered into a framework agreement in connection with the investment. Among other things, the framework agreement builds upon the existing relationship between Tesla and xAI by providing a framework for evaluating potential AI collaborations between the companies. Together, the investment and the related framework agreement are intended to enhance Tesla's ability to develop and deploy AI products and services into the physical world at scale. This investment is subject to customary regulatory conditions with the expectation to close in Q1'2026.", - "path": "TSLA-Q4-2025-Update.pdf-->OTHER UPDATES", + "path": "TSLA-Q4-2025-Update.pdf/OTHER UPDATES", "metadata": { "length": 1213, "summary": "", @@ -1552,7 +1552,7 @@ "chunk_id": "79b23d73-b353-5998-a646-93d6739ec465", "type": "text", "content": "", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK", "metadata": { "length": 0, "summary": "", @@ -1569,7 +1569,7 @@ "chunk_id": "8d4c49a6-c531-5fa6-a77d-ad028537fa52", "type": "text", "content": "We are focused on maximum capacity utilization at our factories. Deliveries and deployments will be impacted by aggregate demand for our products, supply chain readiness and allocation decisions between sale to customers or use for our owned and operated fleet.", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Volume", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Volume", "metadata": { "length": 261, "summary": "", @@ -1609,7 +1609,7 @@ "chunk_id": "426e5119-3ef6-5176-9119-fc044eb90be7", "type": "text", "content": "We will manage the businesses such that we ensure a strong balance sheet, maintaining sufficient liquidity to fund our product roadmap, long-term capacity expansion plans – including further vertical integration – and other expenses.", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Cash", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Cash", "metadata": { "length": 233, "summary": "", @@ -1649,7 +1649,7 @@ "chunk_id": "99f54247-8409-5df2-8299-39184863cd09", "type": "text", "content": "While we continue to execute on innovations to reduce the cost of manufacturing and operations, over time, we expect our hardware-related profits to be accompanied by an acceleration of AI, software and fleet-based profits.", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Profit", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Profit", "metadata": { "length": 223, "summary": "", @@ -1686,7 +1686,7 @@ "chunk_id": "29e89187-3dbe-5298-87c3-c38d3cbe6883", "type": "text", "content": "We continue to evolve and augment our product lineup with a focus on cost, scale and future monetization opportunities via services powered by our AI software. We remain focused on growing our sales volumes through a differentiated and efficiently managed product portfolio, which includes leveraging and optimizing our existing production capacity before building new factories and production lines.\nCybercab, Tesla Semi and Megapack 3 are on schedule for volume production starting in 2026. First generation production lines for Optimus are being installed in anticipation of volume production.\nPHOTOS & CHARTS", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Product", "metadata": { "length": 612, "summary": "", @@ -1772,7 +1772,7 @@ "chunk_id": "bbe9d89e-eebd-56b1-b7c0-7eda12dd3377", "type": "text", "content": "Cybercab, Tesla Semi and Megapack 3 are on schedule for volume production starting in 2026. First generation production lines for Optimus are being installed in anticipation of volume production. 0\n[images/image-5-Tesla Model Y Driving.jpg]", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product-->MODEL Y - 2025 BEST IN CLASS EURO NCAP $^{(1)}$ SMALL SUV", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Product/MODEL Y - 2025 BEST IN CLASS EURO NCAP $^{(1)}$ SMALL SUV", "metadata": { "length": 245, "summary": "", @@ -1835,7 +1835,7 @@ "chunk_id": "02f7e756-798f-5675-9dd9-6ce9123f92d4", "type": "text", "content": " 0\n[images/image-6-Red Tesla on Coastal Road.jpg]", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product-->MODEL 3 - 2025 BEST IN CLASS EURO NCAP $^{(1)}$ LARGE FAMILY CAR", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Product/MODEL 3 - 2025 BEST IN CLASS EURO NCAP $^{(1)}$ LARGE FAMILY CAR", "metadata": { "length": 66, "summary": "", @@ -1884,7 +1884,7 @@ "chunk_id": "494118dd-5788-526b-a139-407d1cfb0d20", "type": "text", "content": " 0\n[images/image-7-Tesla Interior Interface.jpg]", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product-->FSD (SUPERVISED) $^{1}$ – V14 OFFERS UNPARALLELED DRIVER ASSISTANCE", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Product/FSD (SUPERVISED) $^{1}$ – V14 OFFERS UNPARALLELED DRIVER ASSISTANCE", "metadata": { "length": 66, "summary": "", @@ -1933,7 +1933,7 @@ "chunk_id": "6ac7765a-4da5-573a-963b-f35ab3796f6f", "type": "text", "content": " 0\n[images/image-8-Tesla Interior.jpg]", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product-->DRIVERLESS ROBOTAXI - TESTING IN AUSTIN", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Product/DRIVERLESS ROBOTAXI - TESTING IN AUSTIN", "metadata": { "length": 66, "summary": "", @@ -2016,7 +2016,7 @@ "chunk_id": "4cfe7ce9-5f93-5816-9505-ba562683ee42", "type": "text", "content": " 0\n[images/image-9-Tesla Cybertruck in Snow.jpg]\n\n\n### DRIVERLESS ROBOTAXI - TESTING IN AUSTIN 0\n[images/image-10-Tesla Semi Trucks.jpg]\n\nTESLA SEMI - MEGACHARGER NETWORK PLANNED SITES FOR 2026\n\n### CYBERCAB - COLD WEATHER TESTING IN ALASKA 0\n[images/image-11-US Lightning Map.jpg]", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product-->CYBERCAB - COLD WEATHER TESTING IN ALASKA", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Product/CYBERCAB - COLD WEATHER TESTING IN ALASKA", "metadata": { "length": 318, "summary": "", @@ -2102,7 +2102,7 @@ "chunk_id": "07535539-a2f9-5b79-ac6b-d41cf785ec88", "type": "text", "content": " 0\n[images/image-12-Tesla Factory Milestone.jpg]", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product-->GIGAFACTORY SHANGHAI - 9 MILLIONTH VEHICLE PRODUCED (GLOBALLY)", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Product/GIGAFACTORY SHANGHAI - 9 MILLIONTH VEHICLE PRODUCED (GLOBALLY)", "metadata": { "length": 67, "summary": "", @@ -2202,7 +2202,7 @@ "chunk_id": "702bff82-1736-584d-956e-e7bf7eac4039", "type": "text", "content": " 0\n[images/image-13-Tesla Factory Milestone.jpg]\n\n\n### GIGAFACTORY SHANGHAI - 9 MILLIONTH VEHICLE PRODUCED (GLOBALLY) 0\n[images/image-14-Vehicle Delivery Trends.jpg]\n\n\n 0\n[images/image-15-Quarterly Cash Flow.jpg]\n\n\n 0\n[images/image-16-Financial Performance Chart.jpg]", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product-->GIGAFACTORY NEVADA - 6 MILLIONTH DRIVE UNIT PRODUCED", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/Product/GIGAFACTORY NEVADA - 6 MILLIONTH DRIVE UNIT PRODUCED", "metadata": { "length": 327, "summary": "", @@ -2319,7 +2319,7 @@ "chunk_id": "5e17588f-ea71-56df-be53-135da93fb3a0", "type": "text", "content": " 0\n[images/image-17-Projected Vehicle Deliveries.jpg]\n\n\n 0\n[images/image-18-Cash Flow Trends.jpg]\n\n\n 0\n[images/image-19-Financial Performance Forecast.jpg]", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->KEY METRICS TRAILING 12 MONTHS (TTM) (Unaudited)", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/KEY METRICS TRAILING 12 MONTHS (TTM) (Unaudited)", "metadata": { "length": 207, "summary": "", @@ -2369,7 +2369,7 @@ "chunk_id": "c82a6fed-f085-5ec3-b10a-0381723fa3c9", "type": "text", "content": "Total quarterly revenue decreased 3% YoY to \\$24.9B. YoY, revenue was impacted by the following items $^{(1)}$ :\n- decrease in vehicle deliveries\n- lower regulatory credit revenue\n+ growth in Energy Generation and Storage\n+ growth in Services and Other\n+ positive FX impact of \\$0.3B $^{1}$\n\\+ growth in other automotive ancillary sales, partly driven by an increase in FSD subscriptions\n\\+ higher vehicle average selling price (ASP) (excl. FX impact $^{1}$ ), inclusive of mix impact", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->KEY METRICS TRAILING 12 MONTHS (TTM) (Unaudited)-->Revenue", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/KEY METRICS TRAILING 12 MONTHS (TTM) (Unaudited)/Revenue", "metadata": { "length": 484, "summary": "", @@ -2428,7 +2428,7 @@ "chunk_id": "946fabf6-000f-557a-9d81-0b8c1a25aafb", "type": "text", "content": "Our quarterly operating income decreased 11% YoY to \\$1.4B, resulting in a 5.7% operating margin. YoY, operating income was primarily impacted by the following items $^{(1)}$ :\n- increase in SBC and Restructuring and Other charges\n- increase in operating expenses (excl. SBC and Restructuring and Other) driven by AI and other R&D projects and SG&A\n- higher average cost per vehicle due to lower fixed cost absorption for certain models and an increase in tariffs\n- decrease in vehicle deliveries\n- lower regulatory credit revenue\n+ higher vehicle average gross profit due to mix and pricing impacts\n+ growth in Energy Generation and Storage gross profit\n+ growth in Services and Other gross profit\n+ growth in other automotive ancillary sales, partly driven by an increase in FSD subscriptions", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->KEY METRICS TRAILING 12 MONTHS (TTM) (Unaudited)-->Profitability", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/KEY METRICS TRAILING 12 MONTHS (TTM) (Unaudited)/Profitability", "metadata": { "length": 794, "summary": "", @@ -2502,7 +2502,7 @@ "chunk_id": "001d4b8a-b264-5cae-86a5-8c419eabbec0", "type": "text", "content": "Quarter-end cash, cash equivalents and investments was \\$44.1B. The sequential increase of \\$2.4B was primarily the result of positive free cash flow.", - "path": "TSLA-Q4-2025-Update.pdf-->OUTLOOK-->KEY METRICS TRAILING 12 MONTHS (TTM) (Unaudited)-->Cash", + "path": "TSLA-Q4-2025-Update.pdf/OUTLOOK/KEY METRICS TRAILING 12 MONTHS (TTM) (Unaudited)/Cash", "metadata": { "length": 150, "summary": "", @@ -2683,7 +2683,7 @@ "chunk_id": "15c18264-8e1e-5db6-b3c9-70cd181d1f39", "type": "text", "content": "STATEMENT OF OPERATIONS\n(Unaudited)\n\n[tables/table-8 Q4 2024-Q4 2025 Rev.html]\n\nBALANCE SHEET\n(Unaudited)\n\n[tables/table-9 Balance Sheet 2024-25.html]\n\nSTATEMENT OF CASH FLOWS\n(Unaudited)\n\n[tables/table-10 Cash Flow Q4-25.html]\n\nRECONCILIATION OF GAAP TO NON-GAAP FINANCIAL INFORMATION (Unaudited)\n\n[tables/table-11 Q4 2024-Q4 2025.html]\n\nRECONCILIATION OF GAAP TO NON-GAAP FINANCIAL INFORMATION\n(Unaudited)\n\n[tables/table-12 Financial Metrics 2021-25.html]\n\nRECONCILIATION OF GAAP TO NON-GAAP FINANCIAL INFORMATION\n(Unaudited)\n\n[tables/table-13 Financial Data 2022-25.html]\n\n\n[tables/table-14 Financial Metrics 2023-25.html]\n\nTTM = Trailing twelve months\n(1) Beginning in Q1'25, Capital expenditures is presented inclusive of purchases of energy generation and storage systems and all prior periods have been adjusted.\n(2) As a result of the adoption of the new crypto assets standard, the previously reported quarterly periods in 2024 have been recast.\n(3) Beginning in Q1'25, Adjusted EBITDA (non-GAAP) is presented net of digital assets gains and losses and all prior periods have been adjusted.", - "path": "TSLA-Q4-2025-Update.pdf-->FINANCIAL STATEMENTS", + "path": "TSLA-Q4-2025-Update.pdf/FINANCIAL STATEMENTS", "metadata": { "length": 1418, "summary": "", @@ -2816,7 +2816,7 @@ "chunk_id": "31babb3c-1abf-5f64-90c7-e8e8a5384bbd", "type": "text", "content": "Tesla will provide a live webcast of its fourth quarter 2025 financial results conference call beginning at 4:30 p.m. CT on January 28, 2026 at ir.tesla.com. This webcast will also be available for replay for approximately one year thereafter.", - "path": "TSLA-Q4-2025-Update.pdf-->WEBCAST INFORMATION", + "path": "TSLA-Q4-2025-Update.pdf/WEBCAST INFORMATION", "metadata": { "length": 243, "summary": "", @@ -2857,7 +2857,7 @@ "chunk_id": "19e3f236-d15a-5b2a-8589-43d28c5316fe", "type": "text", "content": "When used in this update, certain terms have the following meanings. Our vehicle deliveries include only vehicles that have been transferred to end customers with all paperwork correctly completed. Our energy product deployment volume includes both customer units when installed and equipment sales at time of delivery. \"Net income attributable to common stockholders (non-GAAP)\" is equal to (i) net income attributable to common stockholders before (ii)(a) stock-based compensation expense, net of tax, (b) digital assets (gain) loss, net of tax and (c) release of valuation allowance on deferred tax assets. \"Adjusted EBITDA (non-GAAP)\" is equal to (i) net income attributable to common stockholders before (ii)(a) interest expense, (b) provision for (benefit from) income taxes, (c) depreciation, amortization and impairment, (d) stock-based compensation expense and (e) digital assets loss (gain), net. \"Free cash flow\" is operating cash flow less capital expenditures. Average cost per vehicle is cost of automotive sales divided by new vehicle deliveries (excluding operating leases). \"Days sales outstanding\" is equal to (i) average accounts receivable, net for the period divided by (ii) total revenues and multiplied by (iii) the number of days in the period. \"Days payable outstanding\" is equal to (i) average accounts payable for the period divided by (ii) total cost of revenues and multiplied by (iii) the number of days in the period. \"Days of supply\" is calculated by dividing new car ending inventory by the relevant period's deliveries and using trading days. Constant currency impacts are calculated by comparing actuals against current results converted into USD using average exchange rates from the prior period.", - "path": "TSLA-Q4-2025-Update.pdf-->CERTAIN TERMS", + "path": "TSLA-Q4-2025-Update.pdf/CERTAIN TERMS", "metadata": { "length": 1733, "summary": "This passage defines key financial and operational terms used in a specific update. It clarifies that vehicle deliveries refer to units transferred to end customers with completed paperwork, while energy product deployment includes installed customer units and equipment sales. The text details non-GAAP measures: 'Net income attributable to common stockholders' adjusts for stock-based compensation, digital asset gains/losses, and valuation allowances; 'Adjusted EBITDA' excludes interest, taxes, depreciation, amortization, impairment, stock-based compensation, and digital asset impacts. 'Free cash flow' is defined as operating cash flow minus capital expenditures. Operational metrics include 'Average cost per vehicle' (automotive sales cost divided by new deliveries excluding leases), 'Days sales outstanding' (average receivables divided by revenue times days in period), 'Days payable outstanding' (average payables divided by cost of revenues times days in period), and 'Days of supply' (ending inventory divided by deliveries using trading days). Finally, constant currency impacts are calculated by comparing actuals against results converted to USD using prior period average exchange rates.", @@ -2982,7 +2982,7 @@ "chunk_id": "332edac0-294c-5e5b-8599-a0e52b25e53a", "type": "text", "content": "Consolidated financial information has been presented in accordance with GAAP as well as on a non-GAAP basis to supplement our consolidated financial results. Our non-GAAP financial measures include non-GAAP net income (loss) attributable to common stockholders, non-GAAP net income (loss) attributable to common stockholders on a diluted per share basis (calculated using weighted average shares for GAAP diluted net income (loss) attributable to common stockholders), Adjusted EBITDA margin, non-GAAP automotive gross margin and free cash flow. These non-GAAP financial measures also facilitate management's internal comparisons to Tesla's historical performance as well as comparisons to the operating results of other companies. Management believes that it is useful to supplement its GAAP financial statements with this non-GAAP information because management uses such information internally for its operating, budgeting and financial planning purposes. Management also believes that presentation of the non-GAAP financial measures provides useful information to our investors regarding our financial condition and results of operations, so that investors can see through the eyes of Tesla management regarding important financial metrics that Tesla uses to run the business and allowing investors to better understand Tesla's performance. Non-GAAP information is not prepared under a comprehensive set of accounting rules and therefore, should only be read in conjunction with financial information reported under U.S. GAAP when understanding Tesla's operating performance. A reconciliation between GAAP and non-GAAP financial information is provided above.", - "path": "TSLA-Q4-2025-Update.pdf-->NON-GAAP FINANCIAL INFORMATION", + "path": "TSLA-Q4-2025-Update.pdf/NON-GAAP FINANCIAL INFORMATION", "metadata": { "length": 1664, "summary": "Tesla presents consolidated financial information under both GAAP and non-GAAP standards to supplement its results. Key non-GAAP measures include net income attributable to common stockholders, diluted per share figures, Adjusted EBITDA margin, automotive gross margin, and free cash flow. These metrics aid internal management comparisons with historical performance and other companies, supporting operating, budgeting, and planning activities. Management believes these non-GAAP figures provide investors with a clearer view of Tesla's financial condition and operational results by reflecting the metrics used to run the business. However, since non-GAAP data is not prepared under comprehensive accounting rules, it should be read alongside U.S. GAAP information for a complete understanding of Tesla's performance. A reconciliation between GAAP and non-GAAP data is provided elsewhere.", @@ -3077,7 +3077,7 @@ "chunk_id": "34263854-525e-5a28-b9fd-a88d4813756e", "type": "text", "content": "Certain statements in this update, including, but not limited to, statements in the “Outlook” section; statements relating to the development, strategy, ramp, production and capacity, demand and market growth, cost, pricing and profitability, investment, deliveries, deployment, availability and other features and improvements and timing of existing and future Tesla products and services and supporting infrastructure; statements regarding operating margin, operating profits, spending and liquidity; and statements regarding expansion, improvements and/or ramp and related timing at our facilities are “forward-looking statements” within the meaning of the Private Securities Litigation Reform Act of 1995. Forward-looking statements are based on assumptions and management’s current expectations, involve certain risks and uncertainties, and are not guarantees. Future results may differ materially from those expressed in any forward-looking statement. The following important factors, without limitation, could cause actual results to differ materially from those in the forward-looking statements: the risk of delays in launching and/or manufacturing our products, services and features cost-effectively; our ability to build and/or grow our products and services, sales, delivery, installation, servicing and charging capabilities and effectively manage this growth; our ability to successfully and timely develop, introduce and scale, as well as our consumers’ demand for, products and services based on artificial intelligence, robotics and automation, electric vehicles, advanced driver assistance systems, and ride-hailing services generally and our vehicles and services specifically; the ability of suppliers to deliver components according to schedules, prices, quality and volumes acceptable to us, and our ability to manage such components effectively; any issues with lithium-ion cells or other components manufactured at our factories; our ability to ramp our factories in accordance with our plans; our ability to procure supply of battery cells, including through our own manufacturing; risks relating to international operations and expansion, including unfavorable and uncertain regulatory, political, economic, tax, tariff, export controls and labor conditions; any failures by Tesla products to perform as expected or if product recalls occur; the risk of product liability claims; competition in the automotive, transportation and energy product and services and robotics markets; our ability to maintain public credibility and confidence in our long-term business prospects; our ability to manage risks relating to our various product financing programs; the status of government and economic incentives for electric vehicles and energy products; our ability to attract, hire and retain key employees and qualified personnel; our ability to maintain the security of our information and production and product systems; our compliance with various regulations and laws applicable to our operations and products, which may evolve from time to time; risks relating to our indebtedness and financing strategies; and adverse foreign exchange movements. More information on potential factors that could affect our financial results is included from time to time in our Securities and Exchange Commission filings and reports, including the risks identified under the section captioned “Risk Factors” in our annual report on Form 10-K filed with the SEC on January 30, 2025 and subsequent quarterly reports on Form 10-Q. Tesla disclaims any obligation to update information contained in these forward-looking statements whether as a result of new information, future events or otherwise.\nTESLA", - "path": "TSLA-Q4-2025-Update.pdf-->FORWARD-LOOKING STATEMENTS", + "path": "TSLA-Q4-2025-Update.pdf/FORWARD-LOOKING STATEMENTS", "metadata": { "length": 3711, "summary": "This passage from Tesla outlines that various statements in the update, particularly regarding future outlooks, product development, production capacity, financial metrics, and facility expansions, constitute forward-looking statements under the Private Securities Litigation Reform Act of 1995. These statements reflect management's current expectations based on assumptions and are subject to risks and uncertainties, meaning actual results may differ materially. The text lists numerous specific risk factors that could impact outcomes, including manufacturing delays, supply chain challenges, regulatory hurdles, competition, product performance issues, employee retention, and foreign exchange fluctuations. Tesla advises investors to consult SEC filings, specifically the Form 10-K filed on January 30, 2025, for detailed risk disclosures. The company explicitly disclaims any obligation to update these forward-looking statements due to new information or future events.", diff --git a/apps/api/app/services/demo/source_catalog.py b/apps/api/app/services/demo/source_catalog.py index 054f48b92..534570abb 100644 --- a/apps/api/app/services/demo/source_catalog.py +++ b/apps/api/app/services/demo/source_catalog.py @@ -69,7 +69,7 @@ class DemoSourceDefinition: ), citations=( DemoCitationDefinition( - section_path=("TSLA-Q4-2025-Update.pdf-->OTHER UPDATES"), + section_path=("TSLA-Q4-2025-Update.pdf/OTHER UPDATES"), description="xAI investment", content=( "On January 16, 2026, Tesla entered into an agreement " @@ -92,7 +92,7 @@ class DemoSourceDefinition: citations=( DemoCitationDefinition( section_path=( - "TSLA-Q4-2025-Update.pdf-->SUMMARY-->" + "TSLA-Q4-2025-Update.pdf/SUMMARY/" "Energy generation and storage" ), description="Storage deployment growth", @@ -115,7 +115,7 @@ class DemoSourceDefinition: ), citations=( DemoCitationDefinition( - section_path=("TSLA-Q4-2025-Update.pdf-->OUTLOOK-->Product"), + section_path=("TSLA-Q4-2025-Update.pdf/OUTLOOK/Product"), description="2026 production plans", content=( "Cybercab, Tesla Semi and Megapack 3 are on schedule " diff --git a/apps/api/app/services/demo/source_projection.py b/apps/api/app/services/demo/source_projection.py index ce86b6a7a..8758067af 100644 --- a/apps/api/app/services/demo/source_projection.py +++ b/apps/api/app/services/demo/source_projection.py @@ -5,6 +5,11 @@ from typing import Any, Protocol from urllib.parse import quote +from shared.services.chunks.path_segments import ( + join_document_path, + split_escaped_document_path, +) + _SOURCE_FILE_EXTENSIONS = ( ".csv", ".doc", @@ -241,22 +246,18 @@ def _publication_path( if not raw: return prefix - if "-->" in raw: - sections = [part.strip() for part in raw.split("-->")[1:] if part.strip()] - return "/".join([prefix, *sections]) if sections else prefix - if raw.startswith("images/") or raw.startswith("tables/"): return f"{prefix}/Assets/{raw}" - parts = [part.strip() for part in raw.split("/") if part.strip()] + parts = split_escaped_document_path(raw) if parts and parts[0] == source.title: - return raw + return join_document_path(parts) if len(parts) >= 2 and parts[0] == "Default_Root": section_parts = parts[2:] if parts[1] == source.title else parts[1:] - return "/".join([prefix, *section_parts]) if section_parts else prefix + return join_document_path([prefix, *section_parts]) if section_parts else prefix if parts and _is_source_file_root(parts[0]): section_parts = parts[1:] - return "/".join([prefix, *section_parts]) if section_parts else prefix + return join_document_path([prefix, *section_parts]) if section_parts else prefix return prefix diff --git a/apps/api/tests/contract/test_chunk_document_path_contract.py b/apps/api/tests/contract/test_chunk_document_path_contract.py index d10a99c46..76943f3fd 100644 --- a/apps/api/tests/contract/test_chunk_document_path_contract.py +++ b/apps/api/tests/contract/test_chunk_document_path_contract.py @@ -89,38 +89,6 @@ def test_should_not_treat_dotted_section_titles_as_legacy_document_files() -> No assert children[0]["path"] == chunk_path -def test_should_read_arrow_delimited_document_paths_as_section_paths() -> None: - chunk_path = "report.pdf-->Intro-->Subsection" - - doc_nav = ZipResultSchemaBuilder().build_doc_nav( - [ - { - "chunk_id": "chunk_arrow_delimited_path", - "type": "text", - "content": "arrow delimited content", - "path": chunk_path, - "metadata": {"summary": "arrow delimited summary"}, - } - ], - "report.pdf", - ) - - sections = cast(list[dict[str, object]], doc_nav["sections"]) - children = cast(list[dict[str, object]], sections[0]["children"]) - - assert ( - section_path_from_chunk_path( - chunk_path, - source_file_name="report.pdf", - ) - == "Intro / Subsection" - ) - assert sections[0]["title"] == "Intro" - assert sections[0]["path"] == "report.pdf/Intro" - assert children[0]["title"] == "Subsection" - assert children[0]["path"] == "report.pdf/Intro/Subsection" - - def test_should_preserve_literal_arrow_text_in_slash_paths() -> None: chunk_path = "report.pdf/Inputs --> Outputs/Details" diff --git a/apps/worker/app/services/connect_builder/summary_builder.py b/apps/worker/app/services/connect_builder/summary_builder.py index 960edb836..f55dec47d 100644 --- a/apps/worker/app/services/connect_builder/summary_builder.py +++ b/apps/worker/app/services/connect_builder/summary_builder.py @@ -1,15 +1,16 @@ """ summary_builder: Bottom-up recursive summarization for document navigation. -Reads doc_nav.json + chunks.json for a file and generates ``summary`` fields -at every intermediate node via LLM aggregation. -The enriched doc_nav.json is written back to disk. +Reads doc_nav.json and in-memory chunks, generates ``summary`` at every +non-leaf via deterministic covers assembly or LLM aggregation (with self_only). Usage (standalone): from app.services.connect_builder.summary_builder import enrich_doc_nav_summaries - enrich_doc_nav_summaries(document_workspace_dir, source_file="report.pdf") + enrich_doc_nav_summaries(document_workspace_dir, source_file="report.pdf", chunks=chunks) """ +from __future__ import annotations + import json import os from typing import Any, Dict, List, Optional, Tuple @@ -20,38 +21,78 @@ # ─── Constants ──────────────────────────────────────────────────────────────── -# Summary output max length for recursive LLM aggregation (≤100 chars per node) +# LLM trigger: sum of child contribution lengths + self_only must exceed this SUMMARY_MAX_LEN = 100 -# Navigation top-summary budget measured in semantic tokens (count_cn_en) +# Deterministic (and top-level LLM output) head…tail token budget +DETERMINISTIC_SUMMARY_HEAD = 100 +DETERMINISTIC_SUMMARY_TAIL = 100 +# Navigation top-summary LLM max_tokens NAVIGATION_TOP_SUMMARY_MAX_TOKENS = 200 -# TODO: revisit this cap after we collect more real-world prompt/token budget data. -NON_LLM_TOP_SUMMARY_MAX_SECTIONS = 20 -NON_LLM_TOP_SUMMARY_MAX_DEPTH = 2 + +SECTION_COVERS_PREFIX = "This section covers: " +DOCUMENT_INCLUDES_PREFIX = "This document includes: " # ─── LLM Interface ─────────────────────────────────────────────────────────── -def _llm_summarize(snippets_text: str, node_name: str, max_tokens: int = 100) -> str: - """ - Call LLM to produce a concise summary from aggregated child snippets. +def _build_scope_payload_text( + *, + node_name: str, + self_only: str, + child_rows: List[Tuple[str, str]], +) -> str: + """Flatten SCOPE fields for language detection / logging.""" + titles = [title for title, _ in child_rows] + covered = "\n".join(f"- [{title}] {contrib}" for title, contrib in child_rows) + return "\n".join( + [ + f"SCOPE_TITLE: {node_name}", + f"self_only: {'yes' if self_only.strip() else 'no'}", + f"children: {', '.join(titles)}", + f"SELF_ONLY_CONTENT:\n{self_only.strip() or '(none)'}", + f"COVERED_NODES:\n{covered or '(none)'}", + ] + ) + + +def _llm_summarize( + *, + node_name: str, + self_only: str, + child_rows: List[Tuple[str, str]], + max_tokens: int = 100, +) -> str: + """Call LLM to produce a concise summary for one scope. - Returns plain text summary, or "" on failure. + Returns plain text summary, or "" on failure / null. """ try: from shared.services.ai.prompt_service import build_prompt, _detect_text_language from shared.services.ai.llm_overrides import get_text_client - # Deterministic language lock — see prompt_service._language_directive - detected_lang = _detect_text_language(snippets_text) + payload_text = _build_scope_payload_text( + node_name=node_name, + self_only=self_only, + child_rows=child_rows, + ) + detected_lang = _detect_text_language(payload_text) + child_titles = [title for title, _ in child_rows] + covered_nodes = "\n".join( + f"- [{title}] {contrib}" for title, contrib in child_rows + ) prompt, temperature, top_p, _prompt_max_tokens = build_prompt( task="file-summary", - texts=snippets_text, + texts=payload_text, query="", paras={ "max_tokens": max_tokens, "node_name": node_name, "lang": detected_lang, + "has_self_only": bool(self_only.strip()), + "child_titles": child_titles, + "self_only_content": self_only.strip() or "(none)", + "covered_nodes": covered_nodes or "(none)", }, ) messages: list[ChatCompletionMessageParam] = [ @@ -78,11 +119,7 @@ def _llm_summarize(snippets_text: str, node_name: str, max_tokens: int = 100) -> return "" - -# -# Uses explicit children arrays: -# [{"title": "Section A", "summary": "...", "children": [{"title": "SubA1", ...}]}] -# ═══════════════════════════════════════════════════════════════════════════════ +# ─── doc_nav I/O ───────────────────────────────────────────────────────────── DOC_NAV_FILENAME = "doc_nav.json" @@ -130,6 +167,113 @@ def ensure_doc_nav_json( return nav_path +# ─── self_only + deterministic assembly ────────────────────────────────────── + + +def build_self_only_lookup( + chunks: List[Dict[str, Any]], + *, + source_file_name: str = "", +) -> Dict[str, str]: + """Map canonical section_path → concatenated exact-path chunk content. + + Exact path only (no descendants): same semantics as hydrate ``self_only``. + """ + from shared.services.retrieval.search.lexical_text import section_path_from_chunk_path + + by_path: Dict[str, List[str]] = {} + for chunk in chunks or []: + if not isinstance(chunk, dict): + continue + raw_path = str(chunk.get("path") or "").strip() + if not raw_path: + continue + section_path = section_path_from_chunk_path( + raw_path, + source_file_name=source_file_name, + ) + if not section_path or section_path == "Root": + continue + content = str(chunk.get("content") or chunk.get("text") or "").strip() + if not content: + metadata = chunk.get("metadata") or {} + if isinstance(metadata, dict): + content = str(metadata.get("summary") or "").strip() + if not content: + continue + by_path.setdefault(section_path, []).append(content) + return {path: "\n".join(parts) for path, parts in by_path.items()} + + +def _node_section_path(node: Dict[str, Any], source_file_name: str) -> str: + from shared.services.retrieval.search.lexical_text import section_path_from_chunk_path + + nav_path = str(node.get("path") or "").strip() + if not nav_path: + return "" + return section_path_from_chunk_path(nav_path, source_file_name=source_file_name) + + +def _deterministic_section_summary( + *, + is_top_level: bool, + self_only: str, + child_titles: List[str], +) -> str: + """covers/includes prefix → self_only → child titles; then head…tail whole string.""" + from shared.utils.text_utils import truncate_content_preview + + prefix = DOCUMENT_INCLUDES_PREFIX if is_top_level else SECTION_COVERS_PREFIX + segments: List[str] = [] + self_text = (self_only or "").strip() + if self_text: + segments.append(self_text) + if child_titles: + segments.append(", ".join(child_titles)) + assembled = prefix + " ".join(segments) if segments else prefix.rstrip() + return truncate_content_preview( + assembled, + head=DETERMINISTIC_SUMMARY_HEAD, + tail=DETERMINISTIC_SUMMARY_TAIL, + ) + + +def _child_title_list( + children: List[Dict[str, Any]], + *, + is_top_level: bool, +) -> List[str]: + """All direct child titles (empty titles omitted); top-level skips 'root'.""" + titles: List[str] = [] + for child in children: + title = str(child.get("title") or "").strip() + if not title: + continue + if is_top_level and title.lower() == "root": + continue + titles.append(title) + return titles + + +def _child_contribution_rows( + children: List[Dict[str, Any]], + *, + is_top_level: bool, +) -> List[Tuple[str, str]]: + """(title, contribution) where contribution = summary or title.""" + rows: List[Tuple[str, str]] = [] + for child in children: + title = str(child.get("title") or "").strip() + if is_top_level and title.lower() == "root": + continue + summary = str(child.get("summary") or "").strip() + contrib = summary or title + if not title and not contrib: + continue + rows.append((title or contrib, contrib)) + return rows + + # ─── Recursive summarization on doc_nav sections ───────────────────────────── @@ -137,92 +281,86 @@ def _recursive_summarize_nav( node: Dict[str, Any], use_llm: bool = True, is_top_level: bool = False, + *, + self_only_lookup: Optional[Dict[str, str]] = None, + source_file_name: str = "", ) -> str: """Bottom-up recursive summarization on a doc_nav section node. - Operates on the children-array tree structure of doc_nav.json. - - For each node: - - Leaf (children==[]) → keep existing summary (set during ZIP creation). - - Non-leaf → recursively summarize children, then aggregate. - - Writes summary in-place into ``node["summary"]``. - Returns the summary string. + - Leaf → keep existing summary. + - Non-leaf → recurse children, then: + - use_llm=False → always deterministic covers (titles ± self_only); ignore 100 + - use_llm=True → LLM only if sum(contrib lens)+len(self_only) > SUMMARY_MAX_LEN; + otherwise same deterministic covers path """ - children = node.get("children", []) - title = node.get("title", "") + children = list(node.get("children") or []) + title = str(node.get("title") or "") if not children: - # Leaf node — keep existing summary - existing = (node.get("summary") or "").strip() + existing = str(node.get("summary") or "").strip() if existing: node["summary"] = existing - return node.get("summary", "") + return str(node.get("summary") or "") - # Recurse into children - child_summaries: List[Tuple[str, str]] = [] + lookup = self_only_lookup or {} for child in children: - child_summary = _recursive_summarize_nav(child, use_llm, is_top_level=False) - if child_summary: - child_summaries.append((child.get("title", ""), child_summary)) - - if not child_summaries: - return node.get("summary", "") - - # Aggregate child summaries without hard truncation - aggregated_parts = [] - for name, summary in child_summaries: - aggregated_parts.append(f"[{name}] {summary}") - - aggregated_text = "\n".join(aggregated_parts) + _recursive_summarize_nav( + child, + use_llm=use_llm, + is_top_level=False, + self_only_lookup=lookup, + source_file_name=source_file_name, + ) - max_len = NAVIGATION_TOP_SUMMARY_MAX_TOKENS if is_top_level else SUMMARY_MAX_LEN + section_path = _node_section_path(node, source_file_name) + self_only = "" + if section_path and section_path != "Root": + self_only = str(lookup.get(section_path) or "") + if self_only.strip(): + node["self_summary"] = self_only.strip() + elif "self_summary" in node: + node.pop("self_summary", None) + + child_titles = _child_title_list(children, is_top_level=is_top_level) + child_rows = _child_contribution_rows(children, is_top_level=is_top_level) + contrib_len = sum(len(contrib) for _, contrib in child_rows) + len(self_only) + + deterministic = _deterministic_section_summary( + is_top_level=is_top_level, + self_only=self_only, + child_titles=child_titles, + ) - if len(child_summaries) <= 1 and not is_top_level: - result = child_summaries[0][1] + # use_llm=False → always title/covers path; 100-threshold only applies when LLM is on + if not use_llm: + result = deterministic + elif contrib_len > SUMMARY_MAX_LEN: + max_tokens = ( + NAVIGATION_TOP_SUMMARY_MAX_TOKENS if is_top_level else SUMMARY_MAX_LEN + ) + result = _llm_summarize( + node_name=title, + self_only=self_only, + child_rows=child_rows, + max_tokens=max_tokens, + ) else: - if is_top_level and not use_llm: - titles = [name for name, _ in child_summaries if name.lower() != "root"] - else: - titles = [name for name, _ in child_summaries] - - enum_prefix = "This document includes: " if is_top_level else "This section covers: " - title_enum = enum_prefix + ", ".join(titles) - - if not use_llm: - result = title_enum - else: - total_len = sum(len(s) for _, s in child_summaries) - if total_len > SUMMARY_MAX_LEN: - result = _llm_summarize(aggregated_text, title, max_tokens=max_len) - if not result: - result = title_enum - else: - result = title_enum + result = deterministic node["summary"] = result return result def _doc_nav_has_enriched_summaries(doc_nav: Dict[str, Any]) -> bool: - """Check if enrichment has already been run on this doc_nav. + """True iff every non-leaf has a non-empty summary.""" - Aligned with original enrichment logic: - In doc_nav.json, leaf nodes already have summary from ZIP creation, - so only non-leaf (parent) summaries are set by enrichment. - We recursively check that ALL non-leaf nodes across all depths - have a non-empty summary — if any is missing, enrichment is incomplete. - """ def _check_sections(sections: List[Dict[str, Any]]) -> bool: - """Returns True if all non-leaf nodes in sections have summaries.""" for section in sections: children = section.get("children", []) if not children: continue - # This is a non-leaf node — must have summary from enrichment if not section.get("summary"): return False - # Recurse into children to check deeper non-leaf nodes if not _check_sections(children): return False return True @@ -230,7 +368,6 @@ def _check_sections(sections: List[Dict[str, Any]]) -> bool: sections = doc_nav.get("sections", []) if not sections: return False - # Must have at least one non-leaf to be considered enriched has_non_leaf = any(s.get("children") for s in sections) if not has_non_leaf: return False @@ -238,31 +375,29 @@ def _check_sections(sections: List[Dict[str, Any]]) -> bool: def _build_nav_top_summary( - doc_nav: Dict[str, Any], - use_llm: bool = True + doc_nav: Dict[str, Any], + use_llm: bool = True, + *, + self_only_lookup: Optional[Dict[str, str]] = None, + source_file_name: str = "", ) -> str: - """Build navigation-facing top summary from enriched doc_nav.json. - - Strategy: - Treat all sections as children of a virtual Document node and recursively summarize. - """ + """Build navigation-facing top summary from enriched doc_nav.json.""" sections = doc_nav.get("sections", []) if not sections: return "" - + file_name = source_file_name or str(doc_nav.get("file_name") or "") virtual_doc_node = { "title": "Document Overview", - "children": sections + "children": sections, } - - top_summary = _recursive_summarize_nav( - virtual_doc_node, - use_llm=use_llm, - is_top_level=True + return _recursive_summarize_nav( + virtual_doc_node, + use_llm=use_llm, + is_top_level=True, + self_only_lookup=self_only_lookup, + source_file_name=file_name, ) - - return top_summary def enrich_doc_nav_summaries( @@ -270,6 +405,7 @@ def enrich_doc_nav_summaries( source_file: Optional[str] = None, force: bool = False, use_llm: bool = True, + chunks: Optional[List[Dict[str, Any]]] = None, ) -> Dict[str, str]: """Enrich doc_nav.json with bottom-up recursive summaries. @@ -277,7 +413,8 @@ def enrich_doc_nav_summaries( document_workspace_dir: Absolute path to the temporary document workspace. source_file: If given, only process this file. Otherwise process all. force: If True, regenerate even if summaries already exist. - use_llm: If True, use LLM for multi-child aggregation. + use_llm: If True, use LLM when contribution length sum exceeds threshold. + chunks: In-memory parse chunks for exact-path self_only extraction. Returns: Dict mapping file_name → top-level summary string. @@ -302,9 +439,22 @@ def enrich_doc_nav_summaries( logger.debug(f"No {DOC_NAV_FILENAME} for {file_name}, skipping") continue + source_file_name = str(doc_nav.get("file_name") or file_name) + self_only_lookup = build_self_only_lookup( + list(chunks or []), + source_file_name=source_file_name, + ) + if not force and _doc_nav_has_enriched_summaries(doc_nav): - logger.debug(f"Summaries already exist in {DOC_NAV_FILENAME} for {file_name}, skipping") - results[file_name] = _build_nav_top_summary(doc_nav, use_llm=use_llm) + logger.debug( + f"Summaries already exist in {DOC_NAV_FILENAME} for {file_name}, skipping" + ) + results[file_name] = _build_nav_top_summary( + doc_nav, + use_llm=use_llm, + self_only_lookup=self_only_lookup, + source_file_name=source_file_name, + ) continue logger.info( @@ -312,14 +462,23 @@ def enrich_doc_nav_summaries( f"(mode={mode_label})" ) - # Recursively summarize each top-level section for section in doc_nav.get("sections", []): - _recursive_summarize_nav(section, use_llm=use_llm) + _recursive_summarize_nav( + section, + use_llm=use_llm, + self_only_lookup=self_only_lookup, + source_file_name=source_file_name, + ) _save_doc_nav(file_dir, doc_nav) logger.info(f"✅ doc_nav summaries saved for {file_name}") - top_summary = _build_nav_top_summary(doc_nav, use_llm=use_llm) + top_summary = _build_nav_top_summary( + doc_nav, + use_llm=use_llm, + self_only_lookup=self_only_lookup, + source_file_name=source_file_name, + ) results[file_name] = top_summary return results @@ -329,7 +488,12 @@ def load_nav_top_summary(file_dir: str, file_name: str = "") -> str: """Load doc_nav.json and extract the navigation top summary.""" doc_nav = _load_doc_nav(file_dir) if doc_nav is not None: - return _build_nav_top_summary(doc_nav, use_llm=False) + source_file_name = file_name or str(doc_nav.get("file_name") or "") + return _build_nav_top_summary( + doc_nav, + use_llm=False, + source_file_name=source_file_name, + ) return "" @@ -339,16 +503,6 @@ def build_section_summary_lookup(file_dir: str) -> Dict[str, str]: Keys use the DocumentSection.section_path format produced by ``section_path_from_chunk_path`` (strips the filename prefix, joins remaining parts with ``" / "``). - - Traverses the full section tree at all depths. Used by the publication - pipeline to backfill DocumentSection.summary rows. - - Args: - file_dir: Absolute path to the file-level directory - inside the task-scoped parse workspace. - - Returns: - Dict mapping section_path → summary string (empty dict on any error). """ from shared.services.retrieval.search.lexical_text import section_path_from_chunk_path @@ -375,11 +529,12 @@ def _walk(node: Dict[str, Any]) -> None: for section in doc_nav.get("sections", []): _walk(section) - # Populate Root with the document-level top_summary (tree preview). - # This mirrors GraphNode.properties.top_summary and ensures the - # DocumentSection Root row has a summary for data completeness. if "Root" not in lookup: - top_summary = _build_nav_top_summary(doc_nav, use_llm=False) + top_summary = _build_nav_top_summary( + doc_nav, + use_llm=False, + source_file_name=source_file_name, + ) if top_summary: lookup["Root"] = top_summary diff --git a/apps/worker/app/services/document_ingestion/success_finalization.py b/apps/worker/app/services/document_ingestion/success_finalization.py index d900acc34..29ecd4be3 100644 --- a/apps/worker/app/services/document_ingestion/success_finalization.py +++ b/apps/worker/app/services/document_ingestion/success_finalization.py @@ -157,6 +157,7 @@ def _enrich_document_navigation( document_root_for_enrich, source_file=source_file_name, use_llm=summary_use_llm, + chunks=chunks, ) section_summaries = build_section_summary_lookup(str(add_dir)) except Exception as exc: diff --git a/apps/worker/app/services/document_parser/formats/docx/parser.py b/apps/worker/app/services/document_parser/formats/docx/parser.py index 6ec2d3d87..e195319f1 100755 --- a/apps/worker/app/services/document_parser/formats/docx/parser.py +++ b/apps/worker/app/services/document_parser/formats/docx/parser.py @@ -43,6 +43,7 @@ from shared.core.config import settings from shared.core.exceptions.domain_exceptions import DocxParsingException from shared.core.exceptions.knowhere_exception import KnowhereException +from shared.services.chunks.path_segments import escape_path_segment from shared.utils.chunk_refs import build_chunk_ref, has_chunk_ref from app.services.common.file_loading import load_file_bytes from app.services.common.file_utils import path_handle @@ -670,7 +671,8 @@ def convert_doc2dics( continue # Build tentative path to check for duplicates - tentative_path = doc_name + split_char + key + escaped_doc_name = escape_path_segment(doc_name) + tentative_path = escaped_doc_name + split_char + key # Deduplicate: if path already exists, add suffix if tentative_path in path_counter: @@ -680,7 +682,7 @@ def convert_doc2dics( else: path_counter[tentative_path] = 1 - path_keys.append((doc_name + split_char + key)) + path_keys.append((escaped_doc_name + split_char + key)) bottom_content = joined bottom_tokens = tokenize2stw_remove( [bottom_content], base_llm_paras["stopwords"] @@ -699,11 +701,14 @@ def convert_doc2dics( know_id = gen_str_codes(pure_text) # Use relative_root for path instead of the absolute output directory. path_suffix = key if key.strip() else "" - know_path = ( - split_char.join([relative_root, path_suffix]) - if relative_root and path_suffix - else (relative_root or path_suffix) - ) + if relative_root and path_suffix: + know_path = f"{escape_path_segment(relative_root)}/{path_suffix}" + else: + know_path = ( + escape_path_segment(relative_root) + if relative_root + else path_suffix + ) df_list.append( ParsedRow( content=bottom_content, diff --git a/apps/worker/app/services/document_parser/formats/excel/table_parser.py b/apps/worker/app/services/document_parser/formats/excel/table_parser.py index 8dc10101e..4aa4c5e78 100644 --- a/apps/worker/app/services/document_parser/formats/excel/table_parser.py +++ b/apps/worker/app/services/document_parser/formats/excel/table_parser.py @@ -27,6 +27,7 @@ from shared.core.exceptions.domain_exceptions import TableParsingException from shared.core.exceptions.knowhere_exception import KnowhereException +from shared.services.chunks.path_segments import join_document_path from app.services.common.file_loading import load_file_bytes from app.services.common.file_utils import path_handle from shared.utils.text_utils import tokenize2stw_remove @@ -296,11 +297,13 @@ def _build_excel_table_path( ] if subtable_title: parts.append(_clean_path_segment(subtable_title)) - return "/".join(part for part in parts if part) + return join_document_path(parts) def _clean_path_segment(value: str) -> str: - return str(value).strip().replace("/", "_").replace("\\", "_") + # Keep semantic ``/`` for join_document_path escaping; only neutralize + # filesystem-hostile backslashes in sheet/subtable labels. + return str(value).strip().replace("\\", "_") def _summarize_excel_table( diff --git a/apps/worker/app/services/document_parser/formats/markdown/parse_state.py b/apps/worker/app/services/document_parser/formats/markdown/parse_state.py index 8afab023c..a8d709d38 100644 --- a/apps/worker/app/services/document_parser/formats/markdown/parse_state.py +++ b/apps/worker/app/services/document_parser/formats/markdown/parse_state.py @@ -13,6 +13,7 @@ TextDeferredSummaryTask, ) from app.services.document_parser.support.parser_rows import ParsedRow, ParsedRowsBuilder +from shared.services.chunks.path_segments import escape_path_segment ParserRowValues = list[str | int] @@ -98,11 +99,7 @@ def enter_heading(self, heading: str, level: int) -> None: if item_level < adjusted_level ] - current_heading = ( - heading.replace(self.split_char, "∕") - if self.split_char in heading - else heading - ) + current_heading = escape_path_segment(heading) tentative_names = [item_heading for item_heading, _ in self.path_stack] tentative_names.append(current_heading) tentative_path_parts = [self.relative_root] if self.relative_root else [] diff --git a/apps/worker/app/services/document_parser/formats/text/parser.py b/apps/worker/app/services/document_parser/formats/text/parser.py index 70ac5ce6b..66ef843c4 100755 --- a/apps/worker/app/services/document_parser/formats/text/parser.py +++ b/apps/worker/app/services/document_parser/formats/text/parser.py @@ -8,6 +8,11 @@ from loguru import logger from shared.core.config import settings +from shared.services.chunks.path_segments import ( + append_document_path, + join_document_path, + split_escaped_document_path, +) from shared.utils.chunk_refs import CHUNK_REF_PATTERN from app.services.common.file_loading import load_file_bytes @@ -102,9 +107,8 @@ def postprocess_leaf_dics( summary_len = ProcessingConstants.POSTPROCESS_SUMMARY_LEN merged_dict = {} - split_char = settings.SPLIT_CHAR or "/" for identifier, d in dict_list: - identifier = split_char.join(identifier) + identifier = join_document_path(identifier) if identifier in merged_dict: merged_dict[identifier][content_key].extend(d[content_key]) @@ -116,7 +120,7 @@ def postprocess_leaf_dics( merged_list = [(identifier, v["content"]) for identifier, v in merged_dict.items()] merge_df = pd.DataFrame(merged_list, columns=["path_identifier", "content_lst"]) - merge_df["path"] = merge_df["path_identifier"].apply(lambda x: x.split(split_char)) + merge_df["path"] = merge_df["path_identifier"].apply(split_escaped_document_path) merge_df = merge_df[["path", "content_lst", "path_identifier"]] # TODO rough dividing of contents (need more smart dividing) @@ -139,16 +143,13 @@ def postprocess_leaf_dics( head = row["path_identifier"] if not head: head = "**Preface**" + head_parts = split_escaped_document_path(head) + leaf_title = head_parts[-1] if head_parts else head for k in range(num): - sub_head = ( - head - + split_char - + head.split(split_char)[-1] - + " part " - + str(k + 1) - ) + part_title = f"{leaf_title} part {k + 1}" + sub_head = append_document_path(head, part_title) df_with_divides.loc[len(df_with_divides)] = { - "path": sub_head.split(split_char), + "path": split_escaped_document_path(sub_head), "content_lst": sublists[k], "path_identifier": sub_head, } diff --git a/apps/worker/app/services/document_parser/structure/heading_tree.py b/apps/worker/app/services/document_parser/structure/heading_tree.py index dcad6024d..05f5919c7 100644 --- a/apps/worker/app/services/document_parser/structure/heading_tree.py +++ b/apps/worker/app/services/document_parser/structure/heading_tree.py @@ -3,6 +3,8 @@ import pandas as pd from loguru import logger +from shared.services.chunks.path_segments import append_document_path + def build_tree_from_dataframe( heading_preds: pd.DataFrame, @@ -33,9 +35,7 @@ def build_tree_from_dataframe( node_to_id[node_key] = row_id parent_dict[tree_node_key] = {} - current_path = ( - f"{parent_path}/{tree_node_key}" if parent_path else tree_node_key - ) + current_path = append_document_path(parent_path, tree_node_key) stack.append((level, parent_dict[tree_node_key], tree_node_key, current_path)) return root, node_to_id, id_to_row @@ -114,9 +114,7 @@ def _extract_headings_from_tree( ) if isinstance(children, dict) and children: - current_path = ( - f"{parent_path}/{tree_node_key}" if parent_path else tree_node_key - ) + current_path = append_document_path(parent_path, tree_node_key) results.extend( _extract_headings_from_tree( children, @@ -150,12 +148,12 @@ def _remove_isolated_nodes_recursive( else: result_dict[heading] = _remove_isolated_nodes_recursive( children, - parent_path=f"{parent_path}/{heading}" if parent_path else heading, + parent_path=append_document_path(parent_path, heading), ) elif isinstance(children, dict) and children: result_dict[heading] = _remove_isolated_nodes_recursive( children, - parent_path=f"{parent_path}/{heading}" if parent_path else heading, + parent_path=append_document_path(parent_path, heading), ) else: result_dict[heading] = children diff --git a/apps/worker/app/services/document_parser/support/path_helpers.py b/apps/worker/app/services/document_parser/support/path_helpers.py index c94cf3101..ad3a89612 100644 --- a/apps/worker/app/services/document_parser/support/path_helpers.py +++ b/apps/worker/app/services/document_parser/support/path_helpers.py @@ -5,6 +5,7 @@ from typing import Any from shared.utils.chunk_refs import extract_chunk_refs +from shared.services.chunks.path_segments import join_document_path from app.services.common.file_utils import path_handle SUMMARY_PATH_MARKERS: tuple[str, ...] = ("summary", "\u6458\u8981\u603b\u7ed3") @@ -56,8 +57,7 @@ def flatten_dic2paths( if isinstance(value, dict) and value: flatten_dic2paths(value, new_path, result) else: - split_char = os.getenv("SPLIT_CHAR", "/") - result.append(split_char.join(new_path)) + result.append(join_document_path(new_path)) return result diff --git a/apps/worker/app/services/page_memory/_serialization.py b/apps/worker/app/services/page_memory/_serialization.py index b09e34202..c32fe72ea 100644 --- a/apps/worker/app/services/page_memory/_serialization.py +++ b/apps/worker/app/services/page_memory/_serialization.py @@ -11,6 +11,7 @@ page_scope_info, sort_skeletons, ) +from shared.services.chunks.path_segments import split_escaped_document_path def write_json(path: Path, value: Any) -> None: @@ -94,7 +95,9 @@ def serialize_skeletons(skeletons: list[Any]) -> list[dict[str, Any]]: def build_hierarchy_tree(skeletons: list[Any]) -> dict[str, Any]: hierarchy: dict[str, Any] = {} for skel in sort_skeletons(skeletons): - parts = str(getattr(skel, "section_path", "") or "").split("/") + parts = split_escaped_document_path( + getattr(skel, "section_path", "") or "" + ) section_parts = parts[1:] if len(parts) > 1 else parts current = hierarchy for part in section_parts: diff --git a/apps/worker/app/services/page_memory/fine_hierarchy.py b/apps/worker/app/services/page_memory/fine_hierarchy.py index 0e4466824..92ea603ee 100644 --- a/apps/worker/app/services/page_memory/fine_hierarchy.py +++ b/apps/worker/app/services/page_memory/fine_hierarchy.py @@ -5,6 +5,10 @@ row semantics are designed for raw text. This module calls a dedicated ``page-memory-hierarchy`` prompt and builds deeper ``SectionSkeleton`` entries under each coarse TOC leaf. + +Boundary trimming is code-side: VLM extracts all outline titles in reading +order; ``_collect_candidates`` then drops titles at/before the coarse start +anchor and at/after the next-leaf end anchor. """ from __future__ import annotations @@ -22,6 +26,7 @@ from shared.services.ai.llm_overrides import get_text_client from shared.services.ai.prompt_service import build_prompt from shared.services.ai.response_process_service import eval_response +from shared.services.chunks.path_segments import append_document_path def refine_fat_leaf_skeletons( @@ -29,6 +34,7 @@ def refine_fat_leaf_skeletons( coarse_skeletons: list[SectionSkeleton], tag_results: list[PageTagResult], fat_leaf_pages: set[int], + next_title_by_path: dict[str, str | None] | None = None, model_name: str | None = None, max_tokens: int = 2000, max_depth: int = 6, @@ -37,8 +43,8 @@ def refine_fat_leaf_skeletons( """Refine coarse TOC leaf skeletons using VLM-observed title candidates. For each coarse skeleton that overlaps fat-leaf pages: - 1. Collect real ``observed_titles`` from that leaf using exclusive end - boundaries. + 1. Collect ``observed_titles`` in reading order and trim by start/end + coarse anchors. 2. Ask the page-memory hierarchy prompt to assign relative levels. 3. Rebuild the nested tree and graft it under the coarse leaf. @@ -47,6 +53,7 @@ def refine_fat_leaf_skeletons( if not fat_leaf_pages: return coarse_skeletons + next_map = next_title_by_path or {} refined: list[SectionSkeleton] = [] for index, skeleton in enumerate(coarse_skeletons): @@ -58,19 +65,20 @@ def refine_fat_leaf_skeletons( refined.append(skeleton) continue - # Collect observed titles from tagged pages within this skeleton + end_title = next_map.get(skeleton.section_path) candidates = _collect_candidates( skeleton=skeleton, exclusive_end=exclusive_end, tag_results=tag_results, fat_leaf_pages=fat_leaf_pages, + start_title=skeleton.title, + end_title=end_title, ) if not candidates: refined.append(skeleton) continue - # Run hierarchy LLM and graft the resulting tree under the coarse leaf deeper = _run_hierarchy_on_candidates( candidates=candidates, skeleton=skeleton, @@ -103,10 +111,9 @@ def compute_fat_leaf_pages( A "fat leaf" is a coarse skeleton (TOC leaf node) whose page range exceeds ``min_pages``. Only these pages get VLM title detection. - Uses exclusive end boundaries: each skeleton's scan range ends at - ``next_skeleton.start_page - 1`` to avoid scanning pages that belong - to the next sibling skeleton (the raw ``end_page`` from the hierarchy - locator uses closed-closed intervals that can overlap at boundaries). + Uses exclusive end boundaries when sibling starts are visible in + ``skeletons``. For single-leaf scopes the closed ``end_page`` is used, + which includes the shared boundary page with the next leaf. """ fat_pages: set[int] = set() for idx, skel in enumerate(skeletons): @@ -117,6 +124,31 @@ def compute_fat_leaf_pages( return fat_pages +def build_next_title_by_path( + skeletons: list[SectionSkeleton], +) -> dict[str, str | None]: + """Map each leaf ``section_path`` to the next leaf's title (by start_page).""" + ordered = sorted( + skeletons, + key=lambda item: ( + int(getattr(item, "start_page", 0) or 0), + str(getattr(item, "section_path", "") or ""), + ), + ) + next_map: dict[str, str | None] = {} + for index, skeleton in enumerate(ordered): + next_title: str | None = None + start_page = int(getattr(skeleton, "start_page", 0) or 0) + for later in ordered[index + 1 :]: + later_start = int(getattr(later, "start_page", 0) or 0) + if later_start > start_page: + title = str(getattr(later, "title", "") or "").strip() + next_title = title or None + break + next_map[str(skeleton.section_path)] = next_title + return next_map + + # ── Internal helpers ───────────────────────────────────────────────── @@ -126,17 +158,21 @@ def _collect_candidates( exclusive_end: int, tag_results: list[PageTagResult], fat_leaf_pages: set[int], + start_title: str, + end_title: str | None, ) -> list[dict[str, Any]]: """Gather VLM-observed titles within a skeleton's page range. - Returns list of ``{id, heading, page, prominence}`` sorted by page then - prominence (strongest first within page). + Preserves page ascending order and within-page VLM reading order, then + trims by coarse start/end anchors: + + - Drop the start anchor and everything before it. + - Drop the end anchor (next leaf title) and everything after it. + + Returns list of ``{id, heading, page, prominence}``. """ - candidates: list[dict[str, Any]] = [] + raw: list[dict[str, Any]] = [] tag_by_page = {t.page_index: t for t in tag_results} - candidate_id = 0 - seen: set[str] = set() - parent_key = _title_key(skeleton.title) for page in range(skeleton.start_page, exclusive_end + 1): if page not in fat_leaf_pages: @@ -144,31 +180,92 @@ def _collect_candidates( tag = tag_by_page.get(page) if tag is None or not tag.observed_titles: continue - - # Sort by prominence descending within page - sorted_titles = sorted( - tag.observed_titles, - key=lambda t: -(t.get("prominence") or 0.5), - ) - for title_entry in sorted_titles: - text = title_entry.get("text", "").strip() + # Keep VLM reading order; do not re-sort by prominence. + for title_entry in tag.observed_titles: + text = str(title_entry.get("text", "") or "").strip() if not text or len(text) < 2: continue key = _title_key(text) - if not key or key == parent_key or key in seen: + if not key: continue - seen.add(key) - candidate_id += 1 - candidates.append({ - "id": candidate_id, + raw.append({ "heading": text, "page": page, "prominence": title_entry.get("prominence"), + "key": key, }) + trimmed = _trim_by_coarse_anchors( + raw, + start_title=start_title, + end_title=end_title, + section_path=skeleton.section_path, + ) + + candidates: list[dict[str, Any]] = [] + seen: set[str] = set() + candidate_id = 0 + for item in trimmed: + key = str(item["key"]) + if key in seen: + continue + seen.add(key) + candidate_id += 1 + candidates.append({ + "id": candidate_id, + "heading": item["heading"], + "page": item["page"], + "prominence": item.get("prominence"), + }) return candidates +def _trim_by_coarse_anchors( + raw: list[dict[str, Any]], + *, + start_title: str, + end_title: str | None, + section_path: str, +) -> list[dict[str, Any]]: + """Hard-trim reading-order titles by start/end coarse anchors.""" + start_key = _title_key(start_title) + end_key = _title_key(end_title) if end_title else "" + + trimmed = list(raw) + if start_key: + start_idx = next( + (i for i, item in enumerate(trimmed) if item.get("key") == start_key), + None, + ) + if start_idx is None: + logger.warning( + "[page_memory.fine_hierarchy] start anchor not found for {}; " + "keeping all candidates before end trim", + section_path, + ) + # Fallback: drop exact parent-title duplicates if present. + trimmed = [item for item in trimmed if item.get("key") != start_key] + else: + trimmed = trimmed[start_idx + 1 :] + + if end_key: + end_idx = next( + (i for i, item in enumerate(trimmed) if item.get("key") == end_key), + None, + ) + if end_idx is None: + logger.warning( + "[page_memory.fine_hierarchy] end anchor {!r} not found for {}; " + "skipping back trim", + end_title, + section_path, + ) + else: + trimmed = trimmed[:end_idx] + + return trimmed + + def _run_hierarchy_on_candidates( *, candidates: list[dict[str, Any]], @@ -285,7 +382,7 @@ def _run_hierarchy_on_candidates( while stack and stack[-1][0] >= rel_level: stack.pop() parent_path = stack[-1][1] if stack else skeleton.section_path - section_path = f"{parent_path}/{heading}" + section_path = append_document_path(parent_path, heading) stack.append((rel_level, section_path)) nodes.append({ "section_path": section_path, @@ -341,7 +438,7 @@ def _exclusive_end(skeletons: list[SectionSkeleton], index: int) -> int: return min(skeleton.end_page, min(later_starts) - 1) -def _title_key(title: str) -> str: +def _title_key(title: str | None) -> str: normalized = re.sub(r"\s+", "", str(title or "")).casefold() normalized = re.sub(r"[^\w\u4e00-\u9fff]+", "", normalized) return normalized diff --git a/apps/worker/app/services/page_memory/memory_service.py b/apps/worker/app/services/page_memory/memory_service.py index f34cc42c0..f050f631e 100644 --- a/apps/worker/app/services/page_memory/memory_service.py +++ b/apps/worker/app/services/page_memory/memory_service.py @@ -220,8 +220,8 @@ def _build_page_dataframe( C1 page_renderer → PageRenderResult[] C2 page_plan → PagePlan[] C3 page_tagger → PageTagResult[] - C3b title detection → observed_titles - C4b fine_hierarchy → refined SectionSkeleton[] + C3b title detection → observed_titles (reading-order, untrimmed) + C4b fine_hierarchy → start/end anchor trim + refined SectionSkeleton[] C5 page_assets → assets anchored to refined hierarchy pages C7 assemble node-granularity DataFrame """ @@ -230,6 +230,7 @@ def _build_page_dataframe( collapse_single_child_chains, extract_section_skeletons, ) + from shared.services.chunks.path_segments import join_document_path anatomy = getattr(profile, "anatomy", None) page_count = max(int(profile.page_count or 0), 0) if page_count <= 0: @@ -269,7 +270,7 @@ def _build_page_dataframe( if not skeletons: skeletons = [ SectionSkeleton( - section_path=f"{filename}/Root", + section_path=join_document_path([filename, "Root"]), level=1, start_page=1, end_page=page_count, @@ -279,6 +280,9 @@ def _build_page_dataframe( ) ] + from app.services.page_memory.fine_hierarchy import build_next_title_by_path + + next_title_by_path = build_next_title_by_path(skeletons) coarse_scopes = _build_hierarchy_scopes( skeletons=skeletons, filename=filename, @@ -340,6 +344,7 @@ def _build_page_dataframe( asset_extraction_enabled=asset_extraction_enabled, trace_recorder=trace_recorder, page_memory_config=page_memory_config, + next_title_by_path=next_title_by_path, ) if scope_concurrency <= 1 or len(coarse_scopes) <= 1: @@ -579,6 +584,7 @@ def _run_hierarchy_scope( asset_max_pages: int, trace_recorder: Any | None, page_memory_config: PageMemoryConfig, + next_title_by_path: dict[str, str | None] | None = None, ) -> _ScopeRunResult: from app.services.page_memory.fine_hierarchy import ( compute_fat_leaf_pages, @@ -651,6 +657,7 @@ def _run_hierarchy_scope( fat_leaf_pages=fat_leaf_pages, budget=None, vlm_model=vlm_model, + scan_direction=page_memory_config.scan_direction, max_concurrent=page_memory_config.title_detection_concurrency, ) _record_trace_stage( @@ -667,6 +674,7 @@ def _run_hierarchy_scope( coarse_skeletons=scope_skeletons, tag_results=title_tags, fat_leaf_pages=fat_leaf_pages, + next_title_by_path=next_title_by_path, model_name=_resolve_hierarchy_model(page_memory_config), max_tokens=page_memory_config.hierarchy_max_tokens, max_depth=page_memory_config.max_heading_depth, diff --git a/apps/worker/app/services/page_memory/page_tagger.py b/apps/worker/app/services/page_memory/page_tagger.py index 4476e6a69..6e2a2aab1 100644 --- a/apps/worker/app/services/page_memory/page_tagger.py +++ b/apps/worker/app/services/page_memory/page_tagger.py @@ -6,10 +6,9 @@ to extract summary + keywords from raw text. For ``skip_tagging`` pages, content is preserved but summary is omitted. -Step 2 of page-memory native hierarchy adds: -- Independent VLM title candidate extraction (``observed_titles``) -- Fat-leaf gating: only pages in TOC leaves with > N pages trigger title detection -- Title extraction uses a dedicated verbatim-only prompt (temp=0, small max_tokens) +Title detection (``tag_page_titles``) extracts outline-level headings in +reading order on fat-leaf pages. Coarse start/end trimming is applied later +in ``fine_hierarchy._collect_candidates``. """ from __future__ import annotations @@ -267,10 +266,14 @@ def tag_page_titles( fat_leaf_pages: set[int], budget: Any | None = None, vlm_model: str | None = None, + scan_direction: str = "top_to_bottom_left_to_right", max_concurrent: int | None = None, ) -> list[PageTagResult]: """Run independent VLM title detection on fat-leaf pages. + Extracts all outline-level headings in reading order. Coarse start/end + trimming happens later in ``fine_hierarchy._collect_candidates``. + Parameters ---------- pages: @@ -284,6 +287,8 @@ def tag_page_titles( Deprecated, ignored. Kept for call-site compatibility. vlm_model: VLM model name; falls back to ``$IMAGE_MODEL``. + scan_direction: + Reading-order wording for the title prompt. max_concurrent: Maximum concurrent title-detection calls. @@ -291,6 +296,7 @@ def tag_page_titles( ------- list[PageTagResult] Updated tag results with ``observed_titles`` populated for fat-leaf pages. + Title lists preserve VLM reading order (no re-sort). """ if not fat_leaf_pages: return tag_results @@ -323,7 +329,11 @@ def _detect_one( page_idx: int, page: PageRenderResult, ) -> tuple[int, list[dict[str, Any]]]: - return page_idx, _tag_vlm_titles(page, model=model) + return page_idx, _tag_vlm_titles( + page, + model=model, + scan_direction=scan_direction, + ) import gevent from gevent.pool import Pool as GeventPool @@ -361,13 +371,17 @@ def _tag_vlm_titles( page: PageRenderResult, *, model: str, + scan_direction: str = "top_to_bottom_left_to_right", ) -> list[dict[str, Any]]: """Send page PNG to VLM with the title-only prompt and parse results.""" prompt, temperature, _top_p, max_tokens = build_prompt( "page-memory-vlm-title", "", "", - paras={"max_tokens": 300}, + paras={ + "max_tokens": 300, + "scan_direction": scan_direction, + }, ) try: diff --git a/apps/worker/app/services/page_memory/skeleton_extractor.py b/apps/worker/app/services/page_memory/skeleton_extractor.py index 4b917cc27..bf6be4446 100644 --- a/apps/worker/app/services/page_memory/skeleton_extractor.py +++ b/apps/worker/app/services/page_memory/skeleton_extractor.py @@ -29,6 +29,10 @@ clean_toc_title, ) from loguru import logger +from shared.services.chunks.path_segments import ( + append_document_path, + join_document_path, +) _FRONT_TOC_REGION_GAP_PAGES = 5 @@ -142,6 +146,15 @@ def extract_section_skeletons( ] nodes = toc_nodes + # TODO(page-memory): locate null-page coarse TOC parents (self-only spans). + # Leaf-only closed-closed ranges currently end at the next leaf start, so + # content under a null-page parent (e.g. "Section A Governing requirements") + # before its first child is attributed to the previous leaf. Candidate fix: + # after calibration, place each null-page coarse node A in + # (last leaf under previous same-level coarse, first leaf under A], then + # bound fine-hierarchy scopes with current_leaf → next_coarse. + # See skill pdf-page-based-track Open TODOs. + # Collapse degenerate single-child intermediate chains before locate. # Rule: only merge a parent with its only child when that child is NOT a # leaf (i.e. the child still has children of its own). This preserves the @@ -275,8 +288,12 @@ def _range_to_skeleton( start_page = _clamp_page(item.start_page, page_count) end_page = _clamp_page(item.end_page, page_count) path_titles = [clean_toc_title(title) or title for title in item.path_titles] - section_path = "/".join([filename, *path_titles]) - parent_path = "/".join([filename, *path_titles[:-1]]) if len(path_titles) > 1 else filename + section_path = join_document_path([filename, *path_titles]) + parent_path = ( + join_document_path([filename, *path_titles[:-1]]) + if len(path_titles) > 1 + else filename + ) evidence = { **item.evidence, "resolver": "hierarchy_locator", @@ -1158,7 +1175,7 @@ def _collapse_node(path: str) -> None: new_children: list[str] = [] for gc_path in grandchild_paths: gc = by_path[gc_path] - new_path = f"{node.section_path}/{gc.title}" + new_path = append_document_path(node.section_path, gc.title) promoted = SectionSkeleton( section_path=new_path, level=gc.level - 1, diff --git a/apps/worker/tests/contract/test_page_memory_fine_hierarchy_contract.py b/apps/worker/tests/contract/test_page_memory_fine_hierarchy_contract.py index 4eac1d249..75eb17b94 100644 --- a/apps/worker/tests/contract/test_page_memory_fine_hierarchy_contract.py +++ b/apps/worker/tests/contract/test_page_memory_fine_hierarchy_contract.py @@ -187,3 +187,55 @@ def test_refine_fat_leaf_skeletons_uses_page_memory_prompt_without_demoting_sibl assert refined[1].end_page == 227 assert refined[2].parent_path.endswith("/2 术语") assert all(item.title != skeleton.title for item in refined) + + +def test_refine_fat_leaf_keeps_slash_in_title_as_single_path_segment( + monkeypatch, +) -> None: + skeleton = SectionSkeleton( + section_path="manual.pdf/Index", + level=1, + start_page=10, + end_page=14, + title="Index", + parent_path="manual.pdf", + ) + tags = [ + PageTagResult( + page_index=10, + observed_titles=[ + {"text": "Index", "prominence": 1.0}, + {"text": "Symbols/Numbers", "prominence": 1.0}, + ], + ), + PageTagResult( + page_index=12, + observed_titles=[{"text": "A entries", "prominence": 1.0}], + ), + ] + + monkeypatch.setattr( + fine_hierarchy, + "get_text_client", + lambda requested_model=None: ( + _FakeClient( + [ + {"id": 1, "level": 1}, + {"id": 2, "level": 1}, + ] + ), + requested_model, + ), + ) + + refined = fine_hierarchy.refine_fat_leaf_skeletons( + coarse_skeletons=[skeleton], + tag_results=tags, + fat_leaf_pages={10, 11, 12, 13, 14}, + model_name="test-model", + ) + + assert [item.title for item in refined] == ["Symbols/Numbers", "A entries"] + assert refined[0].section_path == "manual.pdf/Index/Symbols\u2215Numbers" + assert refined[1].parent_path == "manual.pdf/Index" + assert refined[0].section_path.count("/") == 2 diff --git a/apps/worker/tests/contract/test_page_memory_page_tagger_contract.py b/apps/worker/tests/contract/test_page_memory_page_tagger_contract.py index 8eeba6360..22b5104e5 100644 --- a/apps/worker/tests/contract/test_page_memory_page_tagger_contract.py +++ b/apps/worker/tests/contract/test_page_memory_page_tagger_contract.py @@ -32,6 +32,7 @@ def _fake_tag_vlm_titles( page: PageRenderResult, *, model: str, + scan_direction: str = "top_to_bottom_left_to_right", ) -> list[dict[str, object]]: gevent.sleep(0.01 * (4 - page.page_index)) return [{"text": f"title-{page.page_index}", "prominence": 0.8}] @@ -82,6 +83,7 @@ def _fake_tag_vlm_titles( page: PageRenderResult, *, model: str, + scan_direction: str = "top_to_bottom_left_to_right", ) -> list[dict[str, object]]: if page.page_index == 2: raise RuntimeError("title detection failed") @@ -123,6 +125,7 @@ def _fake_tag_vlm_titles( page: PageRenderResult, *, model: str, + scan_direction: str = "top_to_bottom_left_to_right", ) -> list[dict[str, object]]: raise UnavailableException( internal_message="capacity busy", diff --git a/apps/worker/tests/unit/test_summary_builder.py b/apps/worker/tests/unit/test_summary_builder.py new file mode 100644 index 000000000..c0685ef2a --- /dev/null +++ b/apps/worker/tests/unit/test_summary_builder.py @@ -0,0 +1,233 @@ +"""Unit tests for bottom-up doc_nav summary enrichment.""" + +from __future__ import annotations + +from typing import Any, Dict, List + +import pytest + +from app.services.connect_builder.summary_builder import ( + SUMMARY_MAX_LEN, + _deterministic_section_summary, + _llm_summarize, + _recursive_summarize_nav, + build_self_only_lookup, +) + + +def _leaf(title: str, summary: str = "", path: str = "") -> Dict[str, Any]: + node: Dict[str, Any] = {"title": title, "summary": summary, "children": []} + if path: + node["path"] = path + return node + + +def _parent( + title: str, + children: List[Dict[str, Any]], + *, + path: str = "", + summary: str = "", +) -> Dict[str, Any]: + node: Dict[str, Any] = { + "title": title, + "summary": summary, + "children": children, + } + if path: + node["path"] = path + return node + + +class TestDeterministicAssembly: + def test_order_covers_self_only_then_titles(self) -> None: + text = _deterministic_section_summary( + is_top_level=False, + self_only="intro paragraph here", + child_titles=["Alpha", "Beta"], + ) + assert text.startswith("This section covers: ") + assert "intro paragraph here" in text + assert "Alpha, Beta" in text + # self_only before titles + assert text.index("intro paragraph here") < text.index("Alpha, Beta") + + def test_all_child_titles_even_when_summary_empty(self) -> None: + parent = _parent( + "Parent", + [ + _leaf("HasText", summary="body"), + _leaf("NoSummary", summary=""), + ], + path="doc.pdf/Parent", + ) + result = _recursive_summarize_nav( + parent, + use_llm=False, + source_file_name="doc.pdf", + ) + assert "HasText" in result + assert "NoSummary" in result + assert result.startswith("This section covers: ") + + +class TestSelfOnlyLookup: + def test_exact_path_only_excludes_descendants(self) -> None: + chunks = [ + { + "path": "doc.pdf/2.4.4 隐患治理", + "content": "PARENT_INTRO_ONLY", + }, + { + "path": "doc.pdf/2.4.4 隐患治理/清单项A", + "content": "CHILD_BODY_SHOULD_NOT_APPEAR", + }, + ] + lookup = build_self_only_lookup(chunks, source_file_name="doc.pdf") + assert lookup["2.4.4 隐患治理"] == "PARENT_INTRO_ONLY" + assert "CHILD_BODY_SHOULD_NOT_APPEAR" not in lookup["2.4.4 隐患治理"] + assert "2.4.4 隐患治理 / 清单项A" in lookup + + def test_nonleaf_includes_self_only_in_deterministic(self) -> None: + parent = _parent( + "2.4.4 隐患治理", + [ + _leaf("清单项A", summary="a", path="doc.pdf/2.4.4 隐患治理/清单项A"), + _leaf("清单项B", summary="b", path="doc.pdf/2.4.4 隐患治理/清单项B"), + ], + path="doc.pdf/2.4.4 隐患治理", + ) + lookup = build_self_only_lookup( + [{"path": "doc.pdf/2.4.4 隐患治理", "content": "方案包括以下内容:"}], + source_file_name="doc.pdf", + ) + result = _recursive_summarize_nav( + parent, + use_llm=False, + self_only_lookup=lookup, + source_file_name="doc.pdf", + ) + assert "方案包括以下内容:" in result + assert "清单项A" in result + assert "清单项B" in result + assert parent.get("self_summary") == "方案包括以下内容:" + + +class TestLlmTrigger: + def test_short_contrib_skips_llm(self, monkeypatch: pytest.MonkeyPatch) -> None: + called = {"n": 0} + + def _boom(**kwargs: Any) -> str: + called["n"] += 1 + return "SHOULD_NOT_USE" + + monkeypatch.setattr( + "app.services.connect_builder.summary_builder._llm_summarize", + _boom, + ) + parent = _parent( + "P", + [_leaf("A", summary="x"), _leaf("B", summary="y")], + path="doc.pdf/P", + ) + result = _recursive_summarize_nav(parent, use_llm=True, source_file_name="doc.pdf") + assert called["n"] == 0 + assert result.startswith("This section covers: ") + assert "A" in result and "B" in result + + def test_long_contrib_calls_llm_with_title_for_empty_summary( + self, monkeypatch: pytest.MonkeyPatch + ) -> None: + captured: Dict[str, Any] = {} + + def _fake_llm(**kwargs: Any) -> str: + captured.update(kwargs) + return "LLM_SUMMARY" + + monkeypatch.setattr( + "app.services.connect_builder.summary_builder._llm_summarize", + _fake_llm, + ) + long_a = "A" * (SUMMARY_MAX_LEN + 5) + parent = _parent( + "P", + [ + _leaf("HasSummary", summary=long_a), + _leaf("EmptySummary", summary=""), + ], + path="doc.pdf/P", + ) + lookup = {"P": "SELF_ONLY_INTRO"} + # section path from doc.pdf/P is "P" + result = _recursive_summarize_nav( + parent, + use_llm=True, + self_only_lookup=lookup, + source_file_name="doc.pdf", + ) + assert result == "LLM_SUMMARY" + assert captured["self_only"] == "SELF_ONLY_INTRO" + titles = [t for t, _ in captured["child_rows"]] + contribs = {t: c for t, c in captured["child_rows"]} + assert "EmptySummary" in titles + assert contribs["EmptySummary"] == "EmptySummary" + assert contribs["HasSummary"] == long_a + + def test_single_child_with_self_only_does_not_copy_child( + self, monkeypatch: pytest.MonkeyPatch + ) -> None: + monkeypatch.setattr( + "app.services.connect_builder.summary_builder._llm_summarize", + lambda **kwargs: "MERGED", + ) + long_child = "C" * (SUMMARY_MAX_LEN + 1) + parent = _parent( + "P", + [_leaf("OnlyChild", summary=long_child)], + path="doc.pdf/P", + ) + result = _recursive_summarize_nav( + parent, + use_llm=True, + self_only_lookup={"P": "intro"}, + source_file_name="doc.pdf", + ) + assert result == "MERGED" + assert result != long_child + + +class TestPromptPayload: + def test_file_summary_prompt_contains_scope_blocks( + self, monkeypatch: pytest.MonkeyPatch + ) -> None: + captured: Dict[str, Any] = {} + + def _fake_client() -> Any: + class _C: + def chat_completion(self, **kwargs: Any) -> str: + captured["messages"] = kwargs.get("messages") + return "ok" + + return _C() + + monkeypatch.setattr( + "shared.services.ai.openai_compatible_client_sync.get_openai_client", + _fake_client, + ) + # Ensure build_prompt path works + out = _llm_summarize( + node_name="Parent", + self_only="intro text", + child_rows=[("ChildA", "summary A"), ("ChildB", "ChildB")], + max_tokens=100, + ) + assert out == "ok" + user = captured["messages"][1]["content"] + assert "SCOPE_TITLE: Parent" in user + assert "SELF_ONLY_CONTENT:" in user + assert "intro text" in user + assert "COVERED_NODES:" in user + assert "[ChildA] summary A" in user + assert "[ChildB] ChildB" in user + # legacy flat blob prompt removed + assert "You will receive summaries of sub-sections" not in user diff --git a/packages/shared-python/shared/core/config/ai.py b/packages/shared-python/shared/core/config/ai.py index 33d711fd2..b2f612d2c 100644 --- a/packages/shared-python/shared/core/config/ai.py +++ b/packages/shared-python/shared/core/config/ai.py @@ -44,10 +44,6 @@ class AIConfig(BaseModel): "Same alternates as IMAGE_MODEL (e.g. qwen3-vl-32b-instruct)." ), ) - RETRIEVAL_DECOMPOSITION_ENABLED: bool = Field( - default=False, - description="Enable query-decomposition workflow before agentic retrieval.", - ) RETRIEVAL_PLANNER_MODEL: str = Field( default="", description="Reasoning-capable model used by the workflow query planner.", @@ -72,13 +68,6 @@ class AIConfig(BaseModel): default=3, description="Maximum concurrent workflow steps in the same DAG batch.", ) - RETRIEVAL_AGENTIC_INLINE_TABLE_CHAR_LIMIT: int = Field( - default=10000, - description=( - "Maximum table HTML characters inlined into agentic evidence_text. " - "Larger tables are represented by asset URL, path, summary, and keywords." - ), - ) # Runtime LLM controls. LLM_MOCK_ENABLED: bool = Field( diff --git a/packages/shared-python/shared/models/schemas/page_memory_config.py b/packages/shared-python/shared/models/schemas/page_memory_config.py index bbac71ed3..8e01fce3d 100644 --- a/packages/shared-python/shared/models/schemas/page_memory_config.py +++ b/packages/shared-python/shared/models/schemas/page_memory_config.py @@ -34,6 +34,7 @@ class PageMemoryConfig: page_locate_min_emit_depth: int = 2 page_locate_vlm_candidate_page_cap: int = 4 page_locate_full_leaf_sections: bool = False + scan_direction: str = "top_to_bottom_left_to_right" @classmethod def default(cls) -> Self: @@ -147,6 +148,9 @@ def from_mapping(cls, value: object) -> Self: value.get("page_locate_full_leaf_sections"), default.page_locate_full_leaf_sections, ), + scan_direction=_as_str( + value.get("scan_direction"), default.scan_direction + ), ) def to_dict(self) -> dict[str, object]: diff --git a/packages/shared-python/shared/services/ai/prompt_service.py b/packages/shared-python/shared/services/ai/prompt_service.py index 650864be7..35323256d 100755 --- a/packages/shared-python/shared/services/ai/prompt_service.py +++ b/packages/shared-python/shared/services/ai/prompt_service.py @@ -516,24 +516,41 @@ def build_prompt(task, texts, query, **kwargs): elif task == "page-memory-vlm-title": temperature = 0 top_p = 0.01 - max_tokens = kwargs.get("paras", {}).get("max_tokens", 300) - prompt = """\ + paras = kwargs.get("paras", {}) + max_tokens = paras.get("max_tokens", 300) + scan_direction = paras.get("scan_direction", "top_to_bottom_left_to_right") + + if "right_to_left" in scan_direction: + reading_order_upper = "TOP-TO-BOTTOM, RIGHT-TO-LEFT" + column_order = "right to left (i.e. finish the right column before starting the left column)" + else: + reading_order_upper = "TOP-TO-BOTTOM, LEFT-TO-RIGHT" + column_order = "left to right (i.e. finish the left column before starting the right column)" + + prompt = f"""\ You are extracting document-outline-level headings from a PDF page screenshot. - Your goal is to find ONLY the headings that would appear in a Table of Contents. - Most pages will have ZERO such headings — returning an empty list is expected - and correct for the majority of pages. + Your goal is to find ONLY the section headings that structure the document. + If no text on this page qualifies as a section heading, return an empty list. + + READING ORDER: + This page may contain one or more readable columns. + Within each column, read from top to bottom. + Between columns, read from {column_order}. + Return every qualifying heading on this page in that reading order. + Do not skip a heading just because it looks like a known section title; + extract all outline-level headings that appear on the page. Return strict JSON: - { + {{ "titles": [ - { + {{ "text": "", "prominence": <0.0-1.0>, "is_in_table": , "is_in_header_footer": - } + }} ] - } + }} ═══ MANDATORY BOOLEAN FLAGS (CRITICAL) ═══ For EVERY extracted heading, you MUST accurately evaluate these two flags: @@ -557,10 +574,11 @@ def build_prompt(task, texts, query, **kwargs): 3. VISUAL DISTINCTION (supporting): The text is visually set apart from body text — larger font, bold, - centered, or has extra vertical spacing. + centered, extra vertical spacing, or wrapped in a distinctive + background color block. "prominence": 1.0 = most prominent; 0.5 = medium; 0.1 = minor. - Return titles in TOP-TO-BOTTOM order. Text must be EXACT verbatim. + Return titles in {reading_order_upper} order. Text must be EXACT verbatim. ═══ WHAT TO EXCLUDE (critical — read carefully) ═══ @@ -576,8 +594,8 @@ def build_prompt(task, texts, query, **kwargs): organization/document names repeated as running headers, page numbers, book/volume titles used as running headers or footers. - 3. BODY TEXT — Numbered clauses, list items, paragraphs, or running - prose, even if bold or indented. + 3. INLINE TEXT — bullet list items, numbered clauses, or text that continues a paragraph. + These are content items, not section headings, even if bold. 4. CAPTIONS — Figure/table captions, footnotes. @@ -586,9 +604,9 @@ def build_prompt(task, texts, query, **kwargs): with page numbers — those entries are references, not headings. ═══ IMPORTANT ═══ - Many pages consist entirely of tables, body text, or appendix forms. - These pages have NO qualifying headings. Return {"titles": []} for them. - Do NOT force-extract table labels or body text as headings. + Many pages consist entirely of tables, numbered clauses, or appendix forms. + These pages have NO qualifying headings. Return {{"titles": []}} for them. + Do NOT force-extract table labels or numbered items as headings. Return ONLY the JSON object, no markdown fences. """ @@ -600,20 +618,29 @@ def build_prompt(task, texts, query, **kwargs): max_tokens = kwargs["paras"].get("max_tokens", 2000) coarse_context = kwargs["paras"].get("coarse_context", "") coarse_section = f""" -Confirmed coarse parent section: +Confirmed coarse parent section (the scope of this subtree): ''' {coarse_context} ''' +The input candidates already lie strictly INSIDE this coarse parent. The parent's +own title and the next coarse sibling title (if any) have already been removed. +Do NOT restate or invent those coarse endpoint titles. + +Level 1 means the first heading level under this coarse parent. Nest deeper +headings relative to that parent only. + """ if coarse_context else "" prompt = f""" You are constructing a fine-grained document hierarchy for ONE already-bounded PDF segment. The input rows are NOT raw body text. They are clean title -candidates observed directly from page screenshots by a VLM. +candidates observed directly from page screenshots by a VLM, then trimmed to +the interior of one coarse TOC leaf. Your task: - Assign a relative hierarchy level to each real section/table/form heading. -- Level 1 means top-level inside this segment, level 2 is its child, etc. +- Level 1 means top-level under the confirmed coarse parent, level 2 is its + child, etc. - Preserve all legitimate sibling headings. Consecutive same-level headings are normal and MUST NOT be demoted just because no body text appears between rows. - Use page order as reading order. The "prominence" value is visual strength, @@ -1023,6 +1050,14 @@ def build_prompt(task, texts, query, **kwargs): max_tokens = kwargs["paras"].get("max_tokens", 100) node_name = kwargs["paras"].get("node_name", "") lang = kwargs["paras"].get("lang") + has_self_only = bool(kwargs["paras"].get("has_self_only")) + child_titles = kwargs["paras"].get("child_titles") or [] + self_only_content = kwargs["paras"].get("self_only_content", "(none)") + covered_nodes = kwargs["paras"].get("covered_nodes", "(none)") + if isinstance(child_titles, list): + children_repr = ", ".join(str(t) for t in child_titles) if child_titles else "(none)" + else: + children_repr = str(child_titles) or "(none)" lang_directive = _language_directive(lang) lang_rule = ( f"- **LANGUAGE (HARD CONSTRAINT)**: {lang_directive}" @@ -1030,13 +1065,20 @@ def build_prompt(task, texts, query, **kwargs): else "- Your response must be in the SAME LANGUAGE as the input text" ) - prompt = f"""You will receive summaries of sub-sections from a document section called "{node_name}": - ''' - {texts} - ''' + prompt = f"""SCOPE_TITLE: {node_name} + SCOPE_STRUCTURE: + - self_only: {"yes" if has_self_only else "no"} + - children: [{children_repr}] + + SELF_ONLY_CONTENT: + {self_only_content} + + COVERED_NODES: + {covered_nodes} + Your task: {lang_rule} - - Produce ONE concise sentence summarizing ALL sub-sections, no more than {max_tokens} characters + - Produce ONE concise top-level summary of THIS scope (self_only content plus covered nodes), no more than {max_tokens} characters - Output the summary DIRECTLY, no prefixes, no explanations - If the input lacks meaningful text, return exactly: null """ diff --git a/packages/shared-python/shared/services/chunks/dataframe_chunk_converter.py b/packages/shared-python/shared/services/chunks/dataframe_chunk_converter.py index 5e4ec5905..513f8b78b 100644 --- a/packages/shared-python/shared/services/chunks/dataframe_chunk_converter.py +++ b/packages/shared-python/shared/services/chunks/dataframe_chunk_converter.py @@ -337,11 +337,8 @@ def dataframe_to_chunks(df: _ParserDataFrame | None) -> list[Dict[str, JsonValue metadata["file_path"] = embedded_image_path metadata["original_name"] = image_name else: - normalized_path = path.replace("-->", "/") image_name = ( - os.path.basename(normalized_path) - if normalized_path - else f"image_{chunk_id}.jpg" + os.path.basename(path) if path else f"image_{chunk_id}.jpg" ) _, image_extension = os.path.splitext(image_name) if not image_extension: @@ -355,11 +352,8 @@ def dataframe_to_chunks(df: _ParserDataFrame | None) -> list[Dict[str, JsonValue if embedded_table_path: metadata["file_path"] = embedded_table_path else: - normalized_path = path.replace("-->", "/") table_name = ( - os.path.basename(normalized_path) - if normalized_path - else f"table_{chunk_id}.html" + os.path.basename(path) if path else f"table_{chunk_id}.html" ) metadata["file_path"] = f"tables/{table_name}" diff --git a/packages/shared-python/shared/services/chunks/document_path.py b/packages/shared-python/shared/services/chunks/document_path.py index 37baa9943..a06ba231f 100644 --- a/packages/shared-python/shared/services/chunks/document_path.py +++ b/packages/shared-python/shared/services/chunks/document_path.py @@ -4,6 +4,8 @@ import os +from shared.services.chunks.path_segments import unescape_path_segment + _DOCUMENT_FILE_EXTENSIONS = { ".csv", ".atlas", @@ -38,7 +40,7 @@ def split_document_path( source_file_name: str | None = None, ) -> tuple[list[str], list[str]]: """Return ``(root_parts, section_parts)`` for new and legacy chunk paths.""" - parts = _split_path(path, source_file_name=source_file_name) + parts = _split_path(path) if not parts: return [], [] if parts[0] in _MEDIA_ROOT_SEGMENTS and not _is_legacy_namespace_path( @@ -51,63 +53,16 @@ def split_document_path( return parts[: document_index + 1], parts[document_index + 1 :] -def _split_path(path: str | None, *, source_file_name: str | None) -> list[str]: +def _split_path(path: str | None) -> list[str]: raw = str(path or "").strip() - raw_segments = raw.split("/") - source_segment = _normalize_document_file_name(source_file_name) - parts: list[str] = [] - for index, segment in enumerate(raw_segments): - parts.extend( - _split_arrow_document_segment( - segment, - can_split=_can_split_arrow_document_segment( - index=index, - raw_segments=raw_segments, - segment=segment, - source_segment=source_segment, - ), - ) - ) - return parts - - -def _can_split_arrow_document_segment( - *, - index: int, - raw_segments: list[str], - segment: str, - source_segment: str, -) -> bool: - if index == 0: - return True - if index != 1: - return False - - first_segment = raw_segments[0].strip() if raw_segments else "" - if not _is_document_file_segment(first_segment): - return True - - arrow_document_segment = _normalize_document_file_name( - segment.split("-->", 1)[0] - ) - return bool(source_segment and arrow_document_segment == source_segment) - - -def _split_arrow_document_segment(segment: str, *, can_split: bool) -> list[str]: - normalized_segment = segment.strip() - if not normalized_segment: - return [] - if not can_split or "-->" not in normalized_segment: - return [normalized_segment] - - arrow_parts = [ - part.strip() - for part in normalized_segment.split("-->") - if part.strip() + # Chunk paths use ``/`` as the sole hierarchy separator. Titles that contain + # a semantic ``/`` are escaped at construction time (``∕``) and restored here + # so publication/retrieval see one segment per title. + return [ + unescape_path_segment(segment.strip()) + for segment in raw.split("/") + if segment.strip() ] - if arrow_parts and _is_document_file_segment(arrow_parts[0]): - return arrow_parts - return [normalized_segment] def _find_document_index( diff --git a/packages/shared-python/shared/services/chunks/path_segments.py b/packages/shared-python/shared/services/chunks/path_segments.py new file mode 100644 index 000000000..ff4262e44 --- /dev/null +++ b/packages/shared-python/shared/services/chunks/path_segments.py @@ -0,0 +1,82 @@ +"""Escape semantic slashes in document hierarchy path segments. + +Parser chunk paths use ``/`` as a deterministic hierarchy separator +(``file.pdf/Section/Subsection``). Heading titles may themselves contain +``/`` (e.g. ``Symbols/Numbers``). Those semantic slashes must be escaped +before join and unescaped after split so they are not treated as levels. + +The escape character is U+2215 DIVISION SLASH (``∕``), matching the existing +markdown parse-state behavior. +""" + +from __future__ import annotations + +from collections.abc import Sequence + +# Deterministic hierarchy separator used in parser chunk / section paths. +DOCUMENT_PATH_SEP = "/" + +# Replacement for semantic ``/`` inside a single path segment. +# Must not equal DOCUMENT_PATH_SEP and should be rare in natural titles. +ESCAPED_DOCUMENT_PATH_SEP = "\u2215" # ∕ + + +def escape_path_segment(segment: object) -> str: + """Escape separator characters inside one hierarchy title/segment.""" + text = str(segment or "") + if not text: + return "" + # Escape the escape character first so round-trips stay injective. + text = text.replace(ESCAPED_DOCUMENT_PATH_SEP, ESCAPED_DOCUMENT_PATH_SEP * 2) + return text.replace(DOCUMENT_PATH_SEP, ESCAPED_DOCUMENT_PATH_SEP) + + +def unescape_path_segment(segment: object) -> str: + """Restore semantic slashes previously escaped by ``escape_path_segment``.""" + text = str(segment or "") + if not text: + return "" + out: list[str] = [] + i = 0 + esc = ESCAPED_DOCUMENT_PATH_SEP + while i < len(text): + if text.startswith(esc + esc, i): + out.append(esc) + i += 2 + continue + if text.startswith(esc, i): + out.append(DOCUMENT_PATH_SEP) + i += 1 + continue + out.append(text[i]) + i += 1 + return "".join(out) + + +def join_document_path(parts: Sequence[object]) -> str: + """Join hierarchy parts with ``/``, escaping each part first.""" + escaped = [escape_path_segment(part) for part in parts if str(part or "")] + return DOCUMENT_PATH_SEP.join(escaped) + + +def append_document_path(parent_path: object, *segments: object) -> str: + """Append title segments under an already-escaped parent path.""" + parent = str(parent_path or "") + extras = [escape_path_segment(segment) for segment in segments if str(segment or "")] + if not parent: + return DOCUMENT_PATH_SEP.join(extras) + if not extras: + return parent + return DOCUMENT_PATH_SEP.join([parent, *extras]) + + +def split_escaped_document_path(path: object) -> list[str]: + """Split a ``/``-joined hierarchy path and unescape each segment.""" + raw = str(path or "").strip() + if not raw: + return [] + return [ + unescape_path_segment(part) + for part in raw.split(DOCUMENT_PATH_SEP) + if part != "" + ] diff --git a/packages/shared-python/shared/services/retrieval/search/lexical_text.py b/packages/shared-python/shared/services/retrieval/search/lexical_text.py index a177978da..b6647a16d 100644 --- a/packages/shared-python/shared/services/retrieval/search/lexical_text.py +++ b/packages/shared-python/shared/services/retrieval/search/lexical_text.py @@ -8,6 +8,7 @@ from typing import Any, Optional from shared.services.chunks.document_path import split_document_path +from shared.services.chunks.path_segments import split_escaped_document_path from shared.utils.text_utils import tokenize_contents_for_retrieval _SAME_AS_RE = re.compile(r"\[SAME-AS [^\]]+\]") @@ -35,7 +36,7 @@ def split_section_path(path: Optional[str]) -> list[str]: return [] if " / " in raw: return [p.strip() for p in raw.split(" / ") if p.strip()] - return [p.strip() for p in raw.split("/") if p.strip()] + return split_escaped_document_path(raw) def build_lexical_text(value: str) -> str: diff --git a/packages/shared-python/shared/services/storage/zip_doc_navigation.py b/packages/shared-python/shared/services/storage/zip_doc_navigation.py index ff5f52ebf..938d62e58 100644 --- a/packages/shared-python/shared/services/storage/zip_doc_navigation.py +++ b/packages/shared-python/shared/services/storage/zip_doc_navigation.py @@ -5,6 +5,7 @@ from typing import Any from shared.services.chunks.document_path import split_document_path +from shared.services.chunks.path_segments import join_document_path from shared.utils.text_utils import truncate_content_preview @@ -144,7 +145,7 @@ def _build_section_tree( if key not in root_children: root_children[key] = { "title": "Root", - "path": "/".join(root_parts) if root_parts else path, + "path": join_document_path(root_parts) if root_parts else path, "summary": chunk.get("summary", ""), "chunk_count": 0, "_children_map": {}, @@ -161,7 +162,7 @@ def _build_section_tree( if part not in current_level: current_level[part] = { "title": part, - "path": "/".join(full_section_path_parts), + "path": join_document_path(full_section_path_parts), "summary": "", "chunk_count": 0, "_children_map": {}, diff --git a/packages/shared-python/shared/services/storage/zip_result_resources.py b/packages/shared-python/shared/services/storage/zip_result_resources.py index 797b36eed..58d88c667 100644 --- a/packages/shared-python/shared/services/storage/zip_result_resources.py +++ b/packages/shared-python/shared/services/storage/zip_result_resources.py @@ -95,8 +95,7 @@ def _collect_image_files( original_path = chunk.get("path", "") if original_path: _add_candidate(candidate_names, original_path) - normalized_path = original_path.replace("-->", "/") - original_name = os.path.basename(normalized_path) + original_name = os.path.basename(original_path) source_path, matched_name, ext = _resolve_image_source_path( image_files_map, @@ -178,8 +177,7 @@ def _collect_table_files( original_path = chunk.get("path", "") if original_path: - normalized_path = original_path.replace("-->", "/") - original_name = os.path.basename(normalized_path) + original_name = os.path.basename(original_path) _add_candidate(candidate_names, original_path) else: original_name = None @@ -317,7 +315,7 @@ def _normalize_page_citation_ref(value: Any) -> str | None: def _add_candidate(candidates: list[str], value: str | None) -> None: if not value: return - candidate = os.path.basename(str(value).strip().replace("-->", "/")) + candidate = os.path.basename(str(value).strip()) if not candidate: return if candidate.startswith("[") and candidate.endswith("]"): diff --git a/packages/shared-python/shared/tests/test_path_segments.py b/packages/shared-python/shared/tests/test_path_segments.py new file mode 100644 index 000000000..6a22d39ae --- /dev/null +++ b/packages/shared-python/shared/tests/test_path_segments.py @@ -0,0 +1,112 @@ +"""Tests for hierarchy path segment escaping (titles that contain ``/``).""" + +from __future__ import annotations + +import os + +os.environ.setdefault("DATABASE_URL", "postgresql+asyncpg://test:test@localhost/test") +os.environ.setdefault("TMP_PATH", "/tmp/knowhere-test") +os.environ.setdefault("S3_BUCKET_NAME", "test-uploads") +os.environ.setdefault("S3_ACCESS_KEY_ID", "test") +os.environ.setdefault("S3_SECRET_ACCESS_KEY", "test") +os.environ.setdefault("S3_TEMP_PATH", "/tmp") + +from shared.services.chunks.document_path import split_document_path +from shared.services.chunks.path_segments import ( + ESCAPED_DOCUMENT_PATH_SEP, + append_document_path, + escape_path_segment, + join_document_path, + split_escaped_document_path, + unescape_path_segment, +) +from shared.services.retrieval.search.lexical_text import ( + section_path_from_chunk_path, + split_section_path, +) +from shared.services.storage.zip_doc_navigation import ZipDocNavigationBuilder + + +def test_escape_round_trip_for_slash_in_title() -> None: + title = "Symbols/Numbers" + escaped = escape_path_segment(title) + assert ESCAPED_DOCUMENT_PATH_SEP in escaped + assert "/" not in escaped + assert unescape_path_segment(escaped) == title + + +def test_escape_round_trip_for_literal_escape_char() -> None: + title = f"A{ESCAPED_DOCUMENT_PATH_SEP}B/C" + assert unescape_path_segment(escape_path_segment(title)) == title + + +def test_join_and_split_keeps_slash_title_as_one_segment() -> None: + path = join_document_path(["manual.pdf", "Index", "Symbols/Numbers", "Detail"]) + assert path == ( + f"manual.pdf/Index/Symbols{ESCAPED_DOCUMENT_PATH_SEP}Numbers/Detail" + ) + assert split_escaped_document_path(path) == [ + "manual.pdf", + "Index", + "Symbols/Numbers", + "Detail", + ] + + +def test_append_under_escaped_parent() -> None: + parent = join_document_path(["manual.pdf", "Index"]) + child = append_document_path(parent, "Symbols/Numbers") + assert split_escaped_document_path(child)[-1] == "Symbols/Numbers" + assert child.count("/") == 2 + + +def test_split_document_path_unescapes_section_titles() -> None: + chunk_path = join_document_path( + ["manual.pdf", "Index", "Symbols/Numbers"] + ) + root_parts, section_parts = split_document_path( + chunk_path, + source_file_name="manual.pdf", + ) + assert root_parts == ["manual.pdf"] + assert section_parts == ["Index", "Symbols/Numbers"] + + +def test_section_path_from_chunk_path_preserves_slash_title() -> None: + chunk_path = join_document_path( + ["manual.pdf", "Index", "Symbols/Numbers"] + ) + assert ( + section_path_from_chunk_path( + chunk_path, + source_file_name="manual.pdf", + ) + == "Index / Symbols/Numbers" + ) + assert split_section_path("Index / Symbols/Numbers") == [ + "Index", + "Symbols/Numbers", + ] + + +def test_doc_nav_keeps_slash_title_as_single_section() -> None: + chunk_path = join_document_path( + ["manual.pdf", "Index", "Symbols/Numbers"] + ) + doc_nav = ZipDocNavigationBuilder().build_doc_nav( + [ + { + "chunk_id": "chunk_slash_title", + "type": "text", + "content": "index symbols", + "path": chunk_path, + "metadata": {"summary": "symbols"}, + } + ], + "manual.pdf", + ) + sections = doc_nav["sections"] + assert sections[0]["title"] == "Index" + assert sections[0]["children"][0]["title"] == "Symbols/Numbers" + assert sections[0]["children"][0]["path"] == chunk_path + assert sections[0]["children"][0]["children"] == [] diff --git a/page_memory_work_summary_20260624.md b/page_memory_work_summary_20260624.md deleted file mode 100644 index 7f7ed99d5..000000000 --- a/page_memory_work_summary_20260624.md +++ /dev/null @@ -1,437 +0,0 @@ -# Page-Memory 流程改造与测试进度汇总 - -日期:2026-06-24 -文档:`SJSYJ-SC-2024 企业制度汇编(上册).pdf` -当前 debug 目录:`/Users/wuchengke/.knowhere/_debug_parse/SJSYJ-SC-2024 企业制度汇编(上册).pdf/page_memory` - -## 1. 背景问题 - -这轮工作从“检查解析”对话中的异常开始: - -```text -The encrypted content Hand...run. could not be verified. -Reason: Encrypted content could not be decrypted or parsed. -traceid: f63bf5fe74f7bb1859be513dd7e54b22 -``` - -当时 debug 目录里只有: - -```text -page_memory_fine_hierarchy.json -``` - -但没有看到 PAGE-TAG 结果,也无法判断流程到底跑到哪一步。随后确认核心问题不是单个文件缺失,而是 page-memory 的中间产物、trace、scope 组织方式和生产链路不一致,导致 debug 不可控、不透明。 - -## 2. 已确认的目标流程 - -我们对齐后的 page-memory 生产/测试流程是: - -1. 先做 DOC_PROFILE,抽取文档页特征和 TOC hierarchy。 -2. 基于粗 TOC hierarchy 建立 coarse scopes。 -3. 对每个 coarse scope 构建 fine hierarchy。 -4. 只在 fine hierarchy 覆盖的页面范围内做 PAGE-TAG。 -5. 如开启图表资产提取,则同样在 hierarchy 覆盖范围内抽取资产。 -6. 最后汇总生成顶层 `hierarchy.json`、`page_tags.json`、`assets.json`,并继续支持后续 `chunks.json`、`doc_nav.json`、`manifest.json` 生成。 - -测试链路保持生产形状,但可以用 `--fat-only` 只选择最大的粗 scope 进行快速验证。 - -## 3. 已落地的主要改造 - -### 3.1 Trace 统一到顶层 - -已取消 `_doc_agent/trace.json` 的重复落盘,统一写到: - -```text -page_memory/trace.json -``` - -`_doc_agent/` 目前只保留 DOC_PROFILE/TOC 相关原始调试材料: - -```text -_doc_agent/anatomy_map.json -_doc_agent/toc_hierarchies.json -``` - -### 3.2 中间产物精简 - -当前 stop-at fine 的落盘结构已经收敛为: - -```text -page_memory/ - trace.json - hierarchy.json - page_tags.json - _doc_agent/ - anatomy_map.json - toc_hierarchies.json - scopes/ - p225-301/ - scope.json - fine_hierarchy.json - page_tags.json -``` - -不再输出大量重复的临时 JSON,例如旧的 `coarse_tag_scope.json`、`page_memory_fine_hierarchy.json`、`page_tags_pre_hierarchy.json` 等。 - -### 3.3 Scope 目录改为页码范围 - -scope 目录已从 hash/UUID 风格改为页码范围: - -```text -scopes/p225-301/ -``` - -这样 debug 时可以直接看出该 scope 覆盖的页域。 - -### 3.4 `hierarchy.json` 改为可读树形优先 - -顶层 `hierarchy.json` 和 scope 内 `fine_hierarchy.json` 都改成: - -```json -{ - "HIERARCHY": {}, - "nodes": [], - "stats": {} -} -``` - -其中 `HIERARCHY` 是第一字段,形态对齐最终 `manifest.json` 里的 `HIERARCHY` 字段,方便直接肉眼 debug。 - -scope 内的 `fine_hierarchy.json` 额外包含: - -```json -{ - "scope": {} -} -``` - -机器流程仍可继续使用 `nodes`。 - -### 3.5 页域字段精简 - -scope 和 trace 中不再同时记录 `page_ranges` 和完整 `pages` 列表。 - -当前约定: - -```json -{ - "document_page_count": 423, - "page_count": 77, - "page_ranges": [[225, 301]] -} -``` - -保留逐页 `page_index` 的地方仅限实体数据,例如: - -```text -page_tags.json -assets.json -``` - -因为这些文件本身就是逐页/逐资产记录。 - -## 4. 当前测试进度 - -本次从原始 PDF 重新开始测试: - -```text -/Users/wuchengke/Desktop/temp/test_docs/SJSYJ-SC-2024 企业制度汇编(上册).pdf -``` - -执行目标: - -```text -从头跑到最大粗 scope 的 fine hierarchy -``` - -实际命令: - -```bash -uv run python apps/worker/scripts/debug_page_memory.py \ - --file '/Users/wuchengke/Desktop/temp/test_docs/SJSYJ-SC-2024 企业制度汇编(上册).pdf' \ - --fat-only \ - --stop-at fine -``` - -结果:成功,最终状态为: - -```text -stopped_at_fine -``` - -### 4.1 DOC_PROFILE 结果 - -```text -page_count: 423 -toc_pages: [5, 6, 228] -native TOC TitleNode: 58 -native TOC leaf nodes: 44 -``` - -注意:C4 skeleton 阶段日志显示嵌入式目录页 `[228]` 被当前 global TOC 选择逻辑跳过: - -```text -embedded_toc_region_outside_front_cluster -``` - -这属于后续可优化点。 - -### 4.2 C4 Skeleton 定位 - -```text -skeleton_count: 25 -elapsed: 279.36s -``` - -残余定位阶段使用了多轮小窗口渲染 + VLM 确认,耗时较长。 - -### 4.3 最大粗 scope 选择 - -`--fat-only` 本次选中: - -```text -scope_id: p225-301 -page_ranges: [[225, 301]] -page_count: 77 -coarse skeletons before fine: 1 -``` - -### 4.4 Fine hierarchy 结果 - -title detection: - -```text -77 VLM calls -54 titles found -36 pages with observed_titles -``` - -fine hierarchy: - -```text -1 -> 52 skeletons -elapsed: 110.5s -``` - -最终 `hierarchy.json`: - -```json -{ - "stats": { - "node_count": 52, - "page_count": 77, - "page_ranges": [[225, 301]], - "max_depth": 6 - } -} -``` - -顶层 `HIERARCHY` 当前顶级节点: - -```text -安全类 -``` - -## 5. 当前已验证的文件 - -### 5.1 顶层 `hierarchy.json` - -路径: - -```text -page_memory/hierarchy.json -``` - -检查结果: - -```text -top_keys: HIERARCHY, nodes, stats -node_count: 52 -page_count: 77 -page_ranges: [[225, 301]] -max_depth: 6 -``` - -### 5.2 Scope `scope.json` - -路径: - -```text -page_memory/scopes/p225-301/scope.json -``` - -检查结果: - -```json -{ - "scope_id": "p225-301", - "strategy": "fat_only_coarse_scope:refined", - "document_page_count": 423, - "page_count": 77, - "page_ranges": [[225, 301]], - "skeleton_count": 52 -} -``` - -确认:没有 `pages` 长列表。 - -### 5.3 Scope `fine_hierarchy.json` - -路径: - -```text -page_memory/scopes/p225-301/fine_hierarchy.json -``` - -检查结果: - -```text -top_keys: HIERARCHY, nodes, stats, scope -node_count: 52 -page_count: 77 -page_ranges: [[225, 301]] -max_depth: 6 -``` - -### 5.4 `trace.json` - -路径: - -```text -page_memory/trace.json -``` - -检查结果: - -```text -final_status: stopped_at_fine -stage_count: 10 -summary.page_count: 423 -summary.scope_id: p225-301 -``` - -最后几个 stage 的 page_info 均为 compact range: - -```text -C4.coarse_scope page_count=77 page_ranges=[[225, 301]] -C1.render_pages.coarse page_count=77 page_ranges=[[225, 301]] -C2.page_plan.coarse page_count=77 page_ranges=[[225, 301]] -C3b.title_detection page_count=77 page_ranges=[[225, 301]] -C4b.fine_hierarchy fat_leaf.page_count=77 fat_leaf.page_ranges=[[225, 301]] -``` - -## 6. 已跑过的代码检查 - -最近一次相关检查通过: - -```bash -uv run ruff check \ - apps/worker/app/services/page_memory/memory_service.py \ - apps/worker/app/services/page_memory/fine_hierarchy.py \ - apps/worker/scripts/debug_page_memory.py -``` - -```bash -python -m py_compile \ - apps/worker/app/services/page_memory/memory_service.py \ - apps/worker/app/services/page_memory/fine_hierarchy.py \ - apps/worker/scripts/debug_page_memory.py -``` - -```bash -uv run pytest \ - apps/worker/tests/contract/test_page_memory_fine_hierarchy_contract.py \ - apps/worker/tests/contract/test_page_memory_node_assembler_contract.py \ - apps/worker/tests/contract/test_document_agent_budget_contract.py \ - -q -``` - -结果: - -```text -17 passed -``` - -## 7. 当前待讨论/后续优化点 - -### 7.1 `parent_paths` 仍偏长 - -`scope.json` 里目前仍保留 `parent_paths`,虽然不是 page 长列表,但对 debug 阅读来说有些臃肿。 - -可选优化: - -```json -{ - "root_path": "...", - "parent_path_count": 14 -} -``` - -完整 `parent_paths` 可以放入 `trace.json`。 - -### 7.2 C4 skeleton 残余定位耗时较长 - -本次 C4 skeleton 定位耗时约 279 秒,明显比 fine hierarchy 更重。 - -后续可考虑: - -1. 对已定位的 TOC 节点减少 residual VLM。 -2. 对 debug 模式增加更明确的 residual cap。 -3. 对多个 coarse scope 并发定位/处理。 -4. 复用 anatomy + skeleton cache 做快速迭代。 - -### 7.3 嵌入式 TOC 页 228 被跳过 - -本次 DOC_PROFILE 找到目录页 `[5, 6, 228]`,但 skeleton 阶段跳过了嵌入式目录区域 `[228]`。 - -这可能影响后续更细粒度 scope 的粗 hierarchy 完整性,需要单独评估: - -```text -embedded_toc_region_outside_front_cluster -``` - -### 7.4 下一步测试建议 - -建议下一步直接测试: - -```text -fine hierarchy -> PAGE-TAG -``` - -即跑到: - -```text ---stop-at tag -``` - -重点检查: - -1. `scopes/p225-301/page_tags.json` -2. 顶层 `page_tags.json` -3. `trace.json` 中 PAGE-TAG 是否只覆盖 `[[225, 301]]` -4. PAGE-TAG 是否能支撑后续 `chunks.json` 和 `doc_nav.json` - -之后再开启图表资产提取,验证: - -```text -assets.json -scopes/p225-301/assets.json -``` - -## 8. 当前涉及的主要代码文件 - -本轮 page-memory 相关核心改动集中在: - -```text -apps/worker/app/services/page_memory/memory_service.py -apps/worker/app/services/page_memory/fine_hierarchy.py -apps/worker/app/services/page_memory/page_renderer.py -apps/worker/app/services/page_memory/page_assets.py -apps/worker/app/services/page_memory/node_assembler.py -apps/worker/app/services/document_agent/trace.py -apps/worker/app/services/document_agent/visual.py -apps/worker/scripts/debug_page_memory.py -``` - -其中 `apps/worker/scripts/debug_page_memory.py` 是当前测试入口。 -