Skip to content

用 Vue 3 与 pdf.js 在浏览器中实现 PDF 转 Word ​

译注:本文编译自 dev.to 上的一篇技术教程,原文标题为《How to Convert PDF to Word in the Browser with Vue 3 and pdf.js》,作者 sunshey,原文链接见文末。文章介绍了一种纯前端方案,通过 pdf.js 提取 PDF 文本、再用 docx 库生成 Word 文档,全程无需服务端参与。

将 PDF 转换为 Word,需要以结构化的方式提取文本内容并保留基本格式。本文介绍如何用 Vue 3 和 pdf.js 构建一个运行在浏览器中的 PDF 转 Word 转换器。

挑战:从 PDF 中提取文本 ​

PDF 转 Word 的转换涉及以下几个环节:

  1. 解析 PDF 结构以提取文本
  2. 保留基本格式(标题、列表、段落)
  3. 生成结构正确的 .docx 文件
  4. 处理不同的 PDF 版式

技术栈 ​

  • Vue 3,使用 Composition API
  • pdf.js,用于 PDF 解析与文本提取
  • docx,用于生成 Word 文档
  • Vite,用于打包

核心实现 ​

vue
<script setup lang="ts">
import { ref } from 'vue'
import * as pdfjsLib from 'pdfjs-dist'
import { Document, P, Paragraph } from 'docx'

pdfjsLib.GlobalWorkerOptions.workerSrc = 
  `//cdnjs.cloudflare.com/ajax/libs/pdf.js/${pdfjsLib.version}/pdf.worker.min.js`

const file = ref<File | null>(null)
const converting = ref(false)
const result = ref<Uint8Array | null>(null)

async function handleFile(e: Event) {
  const input = e.target as HTMLInputElement
  if (!input.files?.[0]) return
  file.value = input.files[0]
}

async function convertToWord() {
  if (!file.value) return
  converting.value = true

  const arrayBuffer = await file.value.arrayBuffer()
  const pdf = await pdfjsLib.getDocument({ data: arrayBuffer }).promise

  const children = []

  for (let i = 1; i <= pdf.numPages; i++) {
    const page = await pdf.getPage(i)
    const textContent = await page.getTextContent()

    // Extract and structure text
    const pageText = textContent.items.map(item => item.str).join(' ')

    children.push(new P({
      children: [new Paragraph({ text: pageText })]
    }))

    // Add page break (except last page)
    if (i < pdf.numPages) {
      children.push(new P({ children: [] }))
    }
  }

  const doc = new Document({ sections: [{ children }] })
  const buffer = await doc.toBuffer()
  result.value = buffer.buffer as Uint8Array
  converting.value = false
}
</script>

关键实现细节 ​

1. 用 pdf.js 提取文本 ​

使用 pdf.js 获取每一页的文本项:

typescript
const textContent = await page.getTextContent()
const pageText = textContent.items.map(item => item.str).join(' ')

2. 结构保留 ​

按文本项的位置进行分组,以维持段落结构:

typescript
// Group by Y coordinate to detect paragraph breaks
const paragraphs = groupByY(textContent.items)

3. 生成 DOCX ​

使用 docx 库创建规范的 Word 文档:

typescript
const doc = new Document({ sections: [{ children }] })
const buffer = await doc.toBuffer()

局限性 ​

不支持图片 ​

该实现只提取文本,不处理图片。

解决方案:针对图片较多的 PDF,可增加基于 canvas 的图片提取。

格式支持有限 ​

复杂格式(表格、分栏、自定义字体)无法保留。

解决方案:复杂文档建议使用桌面软件处理。

单次提取 ​

文本提取顺序可能与视觉阅读顺序不一致。

解决方案:提取前按位置(先 Y 后 X)对文本项排序。

小结 ​

构建一个浏览器端的 PDF 转 Word 转换器,主要包含以下步骤:

  1. 使用 pdf.js 解析并提取文本
  2. 将文本组织为段落
  3. 用 docx 库生成 DOCX
  4. 提供下载功能

可访问 en.sotool.top/pdf-to-word 体验该方案。

原文链接 ​

https://dev.to/sunshey/how-to-convert-pdf-to-word-in-the-browser-with-vue-3-and-pdfjs-cbh