用 Vue 3 与 pdf.js 在浏览器中实现 PDF 转 Word
译注:本文编译自 dev.to 上的一篇技术教程,原文标题为《How to Convert PDF to Word in the Browser with Vue 3 and pdf.js》,作者 sunshey,原文链接见文末。文章介绍了一种纯前端方案,通过 pdf.js 提取 PDF 文本、再用 docx 库生成 Word 文档,全程无需服务端参与。
将 PDF 转换为 Word,需要以结构化的方式提取文本内容并保留基本格式。本文介绍如何用 Vue 3 和 pdf.js 构建一个运行在浏览器中的 PDF 转 Word 转换器。
挑战:从 PDF 中提取文本
PDF 转 Word 的转换涉及以下几个环节:
- 解析 PDF 结构以提取文本
- 保留基本格式(标题、列表、段落)
- 生成结构正确的 .docx 文件
- 处理不同的 PDF 版式
技术栈
- Vue 3,使用 Composition API
- pdf.js,用于 PDF 解析与文本提取
- docx,用于生成 Word 文档
- Vite,用于打包
核心实现
vue
<script setup lang="ts">
import { ref } from 'vue'
import * as pdfjsLib from 'pdfjs-dist'
import { Document, P, Paragraph } from 'docx'
pdfjsLib.GlobalWorkerOptions.workerSrc =
`//cdnjs.cloudflare.com/ajax/libs/pdf.js/${pdfjsLib.version}/pdf.worker.min.js`
const file = ref<File | null>(null)
const converting = ref(false)
const result = ref<Uint8Array | null>(null)
async function handleFile(e: Event) {
const input = e.target as HTMLInputElement
if (!input.files?.[0]) return
file.value = input.files[0]
}
async function convertToWord() {
if (!file.value) return
converting.value = true
const arrayBuffer = await file.value.arrayBuffer()
const pdf = await pdfjsLib.getDocument({ data: arrayBuffer }).promise
const children = []
for (let i = 1; i <= pdf.numPages; i++) {
const page = await pdf.getPage(i)
const textContent = await page.getTextContent()
// Extract and structure text
const pageText = textContent.items.map(item => item.str).join(' ')
children.push(new P({
children: [new Paragraph({ text: pageText })]
}))
// Add page break (except last page)
if (i < pdf.numPages) {
children.push(new P({ children: [] }))
}
}
const doc = new Document({ sections: [{ children }] })
const buffer = await doc.toBuffer()
result.value = buffer.buffer as Uint8Array
converting.value = false
}
</script>关键实现细节
1. 用 pdf.js 提取文本
使用 pdf.js 获取每一页的文本项:
typescript
const textContent = await page.getTextContent()
const pageText = textContent.items.map(item => item.str).join(' ')2. 结构保留
按文本项的位置进行分组,以维持段落结构:
typescript
// Group by Y coordinate to detect paragraph breaks
const paragraphs = groupByY(textContent.items)3. 生成 DOCX
使用 docx 库创建规范的 Word 文档:
typescript
const doc = new Document({ sections: [{ children }] })
const buffer = await doc.toBuffer()局限性
不支持图片
该实现只提取文本,不处理图片。
解决方案:针对图片较多的 PDF,可增加基于 canvas 的图片提取。
格式支持有限
复杂格式(表格、分栏、自定义字体)无法保留。
解决方案:复杂文档建议使用桌面软件处理。
单次提取
文本提取顺序可能与视觉阅读顺序不一致。
解决方案:提取前按位置(先 Y 后 X)对文本项排序。
小结
构建一个浏览器端的 PDF 转 Word 转换器,主要包含以下步骤:
- 使用 pdf.js 解析并提取文本
- 将文本组织为段落
- 用 docx 库生成 DOCX
- 提供下载功能
可访问 en.sotool.top/pdf-to-word 体验该方案。
原文链接
https://dev.to/sunshey/how-to-convert-pdf-to-word-in-the-browser-with-vue-3-and-pdfjs-cbh