duan19740319 2013-11-04 13:00
浏览 203

使用PHP从PDF文件中提取HTML表格?

I was wondering if it was possible to extract a table of data from a PDF file, into an array or similar so i can import the table data using PHP? I have DomPDF installed to create PDF files, but this does not have options for reading PDF. If i read the PDF file in PHP i get an encoded string:

%PDF-1.5 5 0 obj <>>> endobj 6 0 obj <>stream x��ێ+��W�\`��E���u

Any help would be appreciated.

Adam

  • 写回答

1条回答 默认 最新

  • dpea85385 2016-12-22 17:45
    关注

    This post is pretty old but seems to have a decent amount of views.

    I'm working on a similar project and have had some success with this https://github.com/mgufrone/pdf-to-html . The HTML returns is just a bunch of absolutely positioned p tags, but if the format of your pdfs are consistent you might have some luck working something out to either parse the table or at least get the data you need.

    Just make sure that you have the poppler utilities installed.

    评论

报告相同问题?

悬赏问题

  • ¥15 基于卷积神经网络的声纹识别
  • ¥15 Python中的request,如何使用ssr节点,通过代理requests网页。本人在泰国,需要用大陆ip才能玩网页游戏,合法合规。
  • ¥100 为什么这个恒流源电路不能恒流?
  • ¥15 有偿求跨组件数据流路径图
  • ¥15 写一个方法checkPerson,入参实体类Person,出参布尔值
  • ¥15 我想咨询一下路面纹理三维点云数据处理的一些问题,上传的坐标文件里是怎么对无序点进行编号的,以及xy坐标在处理的时候是进行整体模型分片处理的吗
  • ¥15 CSAPPattacklab
  • ¥15 一直显示正在等待HID—ISP
  • ¥15 Python turtle 画图
  • ¥15 stm32开发clion时遇到的编译问题