How can I parse word documents ".doc", ".docx" to get all the text using golang?
2条回答 默认 最新
- doumian3780 2016-10-22 20:27关注
You can get some inspiration from those projects:
https://github.com/nguyenthenguyen/docx
https://github.com/opencontrol/doc-templateBasically, DOCX is a Zip file with XMLs in it. All the texts are inside
document.xml
What both project do is remove all XML tags, leaving only text intact. You should see if that approach suits you too.
本回答被题主选为最佳回答 , 对您是否有帮助呢?解决 无用评论 打赏 举报