A developer shared an uncomfortable number this week on Juejin: testing a typical enterprise PDF, traditional RAG (the tech that lets AI retrieve answers from documents) extracted only 336 characters of plain text, while the page's core information—titles, button descriptions, framework-introduction screenshots—sat locked inside bitmap images. Roughly half of what was on the page was invisible to the AI.

This is not an isolated case. Product manuals, operations guides, compliance documents, and technical white papers routinely carry large amounts of core information as screenshots, and most AI PDF readers on the market take a pure-text extraction route—structurally missing information.

What this is

An open-source technical solution: the author uses PyMuPDF as the underlying PDF reader (the industry-standard PDF parsing library), then Alibaba's Qwen-VL multimodal model (a large model that can see images and read text simultaneously) to generate a "searchable text description" for each image, then keyword retrieval (no vector database required—i.e., no need to convert documents into mathematical vectors for similarity matching, saving cost) to recall both text chunks and image chunks together, and hands them all to the LLM to generate answers.

The key design choice is "dual-track retention": images are not just converted into text descriptions (which would drop yet another layer of information). Instead, the original image is preserved for the model to "see" directly, while the text description only serves retrieval recall. This is the essential difference from most RAG solutions on the market.

Industry view

We notice this solution is treated by many developers as "good enough": PyMuPDF is the industry-standard PDF library, calling Alibaba Bailian's API has controllable cost, no vector database is needed to run, and it suits small and mid-sized teams running validation scenarios.

But here's a judgment worth flagging: the author himself labeled it "lightweight," "rapid validation scenarios and small-scale document processing." The solution uses keyword retrieval—recall rate will drop as document volume scales; a single-page Vue site demo cannot prove stability under complex layouts; Qwen-VL's accuracy on architecture diagrams and flowcharts also lacks quantitative evaluation. In other words, this is a "problem-solving approach" worth referencing, not a production-ready product.

Another angle being overlooked: today's consumer-market AI PDF tools—from free tiers to enterprise plans—overwhelmingly take the pure-text route. Users pay for AI and still get blind reading. This is a potential product watershed.

Impact on regular people

  • For enterprise IT: if your internal knowledge base is PDF-heavy, sample a few text-and-image mixed documents and visually compare AI answers with the source. In most cases, you'll find key information in the screenshots got missed.
  • For individual professionals: when using ChatPDF, Kimi, and similar tools to read product manuals, contracts, and technical documents, double-check any key terms hidden in screenshots. Don't take the AI's answer at face value.
  • For the consumer market: text-and-image mixed parsing capability will be the next watershed for AI document tools. The maturity of multimodal models will directly determine who survives.