I built a browser-based Office document viewer library with Claude.
Overview
Using Claude, I created a JavaScript library that can display three Microsoft Office file formats whose specifications have been standardized as Office Open XML (OOXML): DOCX, XLSX, and PPTX.
Background
I had been using GitHub Copilot and ChatGPT for coding since around two years ago, but when it came to so-called “vibe coding,” I had mostly just watched from the sidelines when Cline briefly became popular, without trying it myself.
Around last summer, I had my first opportunity to use Claude as an AI agent other than ChatGPT. My interest in MCP was what prompted me to try it. At the time, after doing some research, I found that Claude was the only free agent that supported connections to a self-hosted MCP server running locally. I do not remember being particularly impressed or disappointed by its capabilities as an agent. What appealed to me was simply that its MCP abstraction layer could serve as a bridge between my own applications and generative AI.
After that, I left it alone again for a while. By autumn, however, I had begun hearing people here and there saying how impressive Claude Code was. Since I already felt a certain affinity for Claude because of my experience with MCP, I decided to pay for it and give it a proper try.
The first thing I built was the web application below, called “Stencil Canvas.” It runs entirely in the user’s browser and transforms an image so that it looks as though it was printed using a stencil-printing process.
https://yukiyokotani.github.io/stencil-canvas/
I will not go into detail here because it is unrelated to the main topic, but I was genuinely amazed by how smoothly Claude fulfilled my requests, including implementing the shaders, without the process ever being derailed by errors.
Seeing that made me want to tackle the main subject of this article: implementing a viewer library for Office files.
I had actually wanted something like this before and had researched the available options. Among open-source Office file viewers, however, most targeted only one of Word, Excel, or PowerPoint, and their rendering quality was generally far from impressive. Commercially licensed libraries offered much better rendering quality, and some even supported editing, but their license fees were naturally substantial. Besides, I did not need anything that feature-rich. If I wanted all of that functionality, I might as well use the official Microsoft 365 applications.
In other words, I could not find anything that was quite right for my needs. All I wanted was a way to render—or preview—Office file binaries with a reasonable degree of fidelity. Nothing more.
During that earlier research, I learned that the specifications for DOCX, XLSX, and PPTX had been standardized and published as OOXML. Having played around with Claude and seen how far generative AI had advanced—reaching the point where it could turn one casual idea after another into reality—asking Claude to implement viewers for three formats with published specifications seemed entirely feasible.
The timing was also convenient because I wanted to catch up on the latest developments around Claude. I therefore upgraded to the Max plan and started working on the project.
What I Built
The result was the following two projects.
The first is a JavaScript/TypeScript library. To prioritize performance, the parser is written in Rust and compiled to WebAssembly. The renderer is implemented in plain TypeScript so that it can be used with any framework.
I paid considerable attention to rendering quality. It is not yet a perfect match for the official Office applications, but I believe it has already reached a very respectable level. Although, to be fair, Claude was the one that implemented it.
The second project is a VS Code extension that uses the first JavaScript library to display Office files in a VS Code Webview.
As an additional optional feature, the extension provides an MCP server with file-reading tools built on the DOCX, XLSX, and PPTX parsers. When you want Claude or GitHub Copilot to use an Office file while coding, an agent without such a tool will typically write a disposable Python script, unzip the file, parse the XML, and read its contents. Having this happen every single time seemed extremely wasteful, so it was something I personally wanted to improve. I think this feature will be particularly useful to a certain group of people.
https://www.npmjs.com/package/@silurus/ooxml
https://marketplace.visualstudio.com/items?itemName=silurus.office-open-xml-viewer
Memories from the Implementation Process
The implementation was done entirely by Claude, and I did not write a single line of code myself. However, I did make several contributions in areas that, at least for now, could only be handled by me as a human, so I would like to record them here. That said, it is possible that an agent could have overcome all of these challenges as well if I had set things up properly and granted it sufficient permissions.
Providing Ground-Truth Data
This was the most important part, and it was something I only came to understand after actually asking Claude to do the work.
The OOXML specifications are not complete when it comes to reproducing how Office applications display documents. I learned this from Claude’s responses: some decisions are left to the implementation. For those cases, I had to provide ground truth on an ad hoc basis—specifically, examples of how Word, Excel, and PowerPoint actually rendered the content.
I began by preparing several Office files as demo data. For the DOCX and PPTX files, I painstakingly exported their pages or slides as images and had Claude set up visual regression testing, or VRT, based on those images.
In the end, however, this approach was almost entirely unsuccessful.
VRT is useful for detecting layout regressions by comparing the current output with a previous snapshot. When building a layout from scratch, however, the pixel-level differences from the ground truth are simply too large. It is also extremely difficult for a machine to interpret what is wrong—for example, whether an element is missing or merely looks different. As a result, this approach did virtually nothing to improve rendering accuracy.
I therefore changed course and exported the DOCX and PPTX files as PDFs, using those PDFs as the ground truth instead.
Claude appeared to parse the reference PDFs as needed using tools such as PDF.js. This allowed the agent to determine precisely which elements should be drawn and at which coordinates. From there, it could identify the relevant parts of the OOXML specifications and translate them into an implementation.
This dramatically improved the rendering quality and allowed us to complete the baseline implementation. We are now at the stage of addressing the remaining discrepancies one at a time, with me identifying them visually.
For XLSX files, exporting to PDF causes layout problems because the PDF page model does not align naturally with the spreadsheet model. From the outset, we therefore implemented and improved XLSX rendering by having me visually identify discrepancies one by one. Fortunately, Excel makes it easy to communicate exactly where something is wrong by referring to cell coordinates, so I think improvements proceeded relatively smoothly even with this approach.
I have not abandoned pixel-based VRT altogether. As mentioned above, comparing our output with the ground truth produced by Office did not yield a meaningful metric. Instead, I redefined its role: we now compare the current rendering output with our own previous results and use those comparisons to detect regressions before and after refactoring.
Managing Claude’s Usage Limits
At first, I worked in a single session using Plan mode and switched between Opus and Sonnet. I soon realized, however, that this setup did not come close to using the full session allowance of the Max plan, so I moved to a multi-session workflow.
This significantly improved efficiency while establishing the baseline parsers and renderers for DOCX, XLSX, and PPTX. Eventually, however, the work began to require the kind of manual visual adjustment described above. I became the bottleneck, so I returned to a single-session workflow.
That once again left me unable to use the full session allowance. At the same time, manually approving actions was gradually becoming tiresome, so I now run everything using the automatic mode with Opus. With this setup, even sequential execution on Max (5x) can easily hit both the session and weekly limits. Still, I decided that this was preferable to leaving the allowance unused.
In the end, I was able to experiment with a variety of workflows and develop a practical sense of what can be accomplished with the $100 plan, which was worthwhile in itself.
A Turning Point in History? — Added June 13
Claude Fable 5 became available on June 9, 2026, and I immediately began using it to develop this library.
At first, I simply switched the model to Fable and asked it to conduct a comprehensive review of the existing implementation in a single session. It launched a large number of subagents, all of which apparently ran on Fable as well. As a result, I reached the five-hour session limit in only 15 to 20 minutes.
Previously, I would only approach the limit if I stayed with a single Opus 4.8 session and continuously issued instructions myself, so this came as a shock. GitHub Copilot indicated that Fable consumed twice as many tokens as Opus, but I had not expected its token consumption to be quite this intense.
Starting with the next session, I used Fable for planning and Opus 4.8 for execution. This restored an experience reasonably close to what I was accustomed to.
I do not have the expertise required to fully evaluate Fable’s capabilities. Even so, I caught glimpses of its intelligence in the way it quickly identified problems in the Opus 4.8 implementation—including perfectly valid architectural flaws—and completed the entire assignment without stopping, apparently filling in contextual gaps in my instructions along the way. Admittedly, the implementation had gone through many twists and turns, so some technical debt was inevitable.
Its “eyes” for graphical interfaces, however, were honestly still unimpressive. On one occasion, I reported a Canvas rendering bug using a screenshot and a written explanation. Fable completely misunderstood the issue, changed code that did not need to be changed, and took quite some time to reach the real problem.
If a human had been given the same information—even someone who was not an expert—I am certain they would have understood the problem immediately. That experience reminded me that there is still a long way to go in this area.
After completing that review, I gave Fable its next major task: designing and planning OffscreenCanvas support for the viewer. The work was completed late at night on June 12, at which point I hit the weekly usage limit.
My limit was scheduled to reset at 3:00 p.m. on June 13, so I planned to resume work then. However, on the morning of June 13, the US government suddenly imposed export restrictions on Fable. The reported implication was that Fable would no longer be available outside the United States, although it initially became unavailable to users worldwide.
Fortunately, Fable had already produced a detailed plan for my task on the assumption that Opus would handle the implementation, so I was able to complete that work without any particular problems. Judging from the news, however, it seemed unlikely that I would be able to use Fable for future tasks.
I cannot say that my current design and implementation work is impossible without Fable. Still, once I had used it—and especially after seeing it identify so many issues in the Opus implementation—I began to doubt whether the output produced by Opus was really good enough.
With Fable, I could simply tell myself, “If even Fable cannot do it, then it probably cannot be helped.” That sense of reassurance made life much easier for someone like me, whose role was essentially to pass work along to the model.
I ran tasks with Opus throughout June 13, but somewhere in the back of my mind, I found myself missing Fable.
Fable was subsequently made available again on July 1, with a system that falls back to Opus 4.8 for certain tasks.
The Arrival of GPT-5.6 Sol — Added August 1
OpenAI publicly released GPT-5.6 Sol on July 9, 2026.
Development had continued smoothly with the Fable-and-Opus combination after Fable became available again on July 1. I therefore felt somewhat apprehensive and hesitant about changing my setup, but I cautiously decided to give Sol a try.
I began by using it to fix small bugs in the library. Without being asked, it took the reliable approach of reproducing each problem with test-driven development before resolving it. Both the speed at which it solved problems and the quality of its final output were excellent, giving me the impression that it would be highly capable as an implementation model.
I later paired it with Fable and entrusted it with design tasks as well. It frequently prevailed even in discussions with Fable, confirming that it was an exceptionally capable model.
Then there was the question of hierarchy between the agents. As the work progressed, I wondered whether to use Sol under Fable or Fable under Sol. Because having Fable perform implementation work consumed tokens quickly and rapidly exhausted both the five-hour and weekly limits, I eventually settled on a stable workflow in which Sol handled everything from design through implementation, while Fable served as a reviewer and adviser at key points.
As part of the post-release campaign, the five-hour limit was removed, and Tibo frequently reset the weekly limits. Almost before I knew it, I had shifted toward using Codex, despite having previously depended so heavily on Claude.
The limit resets, announced almost daily through Tibo’s lighthearted posts, felt almost surreal. From a user’s perspective, however, they could not have been more welcome. The “SAINT TIBO / GIVER OF TOKENS RESETER OF LIMITS” meme emerged, and people began pleading for resets on X.
As a result, I stopped using Claude’s Opus 4.8 altogether. Opus 5 was released sometime later, but the current situation is that I have barely used that either.
I could certainly see improvements in Opus 5—although Opus 4.8 was broken, so perhaps “improvement” is not quite the right word. Still, when viewed as an implementation model, Opus 5 offered nothing that made me want to switch away from Sol. Being able to run a frontier model as far as I want is simply extraordinarily comfortable.
One caveat is that, as others have also noted, Sol has a tendency to overengineer—or perhaps to overthink things. Once, I gave it a goal involving a large-scale refactoring effort, and it continued working on it for almost three days straight. That was honestly a bit much.
Growth in Installs, the Arrival of GPT-6 Astra, and More (Sep. 13 Update)
It has been two months, so I thought I’d record where things stand again.
First, both npm installs and GitHub stars have been growing. The package is now at 60,000 weekly installs on npm and 790 stars on GitHub. The growth hasn’t exactly been linear; there were two distinct jumps, both triggered by the project being featured on Hacker News.
https://news.ycombinator.com/item?id=48436863
https://news.ycombinator.com/item?id=49523361
The first time was actually back in June, and the reactions were fairly negative overall. Some comments were quite harsh, including one that said, “It’s 100% hallucinated.”
To be fair, at the time, I myself mostly saw the project as something I had put out into the world to see what would happen. I was well aware that the rendering quality still had plenty of problems, so I couldn’t really blame people for reacting that way. More importantly, I was confident that this was exactly the kind of problem generative AI was well suited to tackling, and I was equally convinced that there was real demand for it. So the criticism didn’t bother me at all.
The second time the project appeared on Hacker News was in September. In the intervening months, a handful of people had started paying attention to the library. Issues and PRs had begun trickling in on GitHub, and those contributions had in turn driven further improvements.
Perhaps because of that, the comments the second time around felt almost like the opposite of June: most of them were positive. As a result, if I remember correctly, weekly npm installs jumped from around 20,000 to 40,000 almost immediately, while the GitHub star count increased by roughly 200.
That exposure also seems to have led to even more people opening issues and submitting PRs. It feels like awareness and adoption are gradually growing. Bug reports about rendering problems that include sample files are especially valuable, and it has been striking to experience firsthand how open-source software gets polished through exactly this kind of feedback.
That’s the story of my own library. Separately, during the same period, I became aware of several other projects pursuing goals similar to mine.
To varying degrees, all of them are trying to build software that handles OOXML through what you might call AI-assisted “vibe coding.” On the one hand, this reinforces just how much demand there is in this area. On the other, each project has a subtly different scope, and I find it interesting how different the results can be even when people are pursuing broadly similar goals with broadly similar tools.
A brief digression: I still think implementing software that handles OOXML with AI turned out to be an exceptionally good problem to choose.
I’ve written about this before, but there are several reasons. The specification is publicly available and largely mature. Microsoft Office provides a reference implementation whose output can serve as the correct answer. And despite that, no open-source implementation capable of fully solving the problem has emerged over several decades.
Web browsers might look superficially similar: their specifications are public, and there are existing browsers that implement them. But web standards are never going to become “finished” in the same sense, and given that excellent browsers such as Chrome and Firefox are already available for free, I personally see little reason to implement my own browser with AI beyond satisfying an engineer’s technical curiosity.
Even if the exercise were intended as a benchmark for generative AI, I think the significance of the task is fundamentally different when there are people who genuinely want the resulting software.
(Of course, I think implementing a browser with AI can have enormous value to the person doing it. I’m not trying to dismiss that kind of project; I’ll add that disclaimer just to be safe.)
Anyway, back to OOXML.
Several generative-AI-driven OOXML projects have now appeared, and their scopes turn out to differ in subtle but important ways.
In my case, I deliberately narrowed the project’s scope to viewing only, in order to limit what I’m responsible for. Within that boundary, I’m trying to push rendering as far as possible toward a reasonable representation of Office documents. I deliberately say “reasonable” rather than “exact,” because there are cases where browser-based rendering and desktop Office fundamentally cannot produce identical results.
As I wrote in the README, once you start supporting editing, the implementation scope expands dramatically. Worse, when dealing with data structures that contain complex dependencies, there is a real risk of corrupting the document—and I certainly don’t trust myself to vibe-code my way through every possible case of that safely.
So I chose to focus exclusively on viewing, where even if there is a bug, the impact can be contained: the user at least has a chance to notice that something looks wrong. Editing was cut from the scope entirely.
(And, as an aside, part of me also thinks that if you need full editing, you might as well just use Microsoft 365.)
What’s interesting is that, as far as I know, every comparable project appearing now includes editing in its scope. I’m impressed by how ambitious that goal is, while at the same time—speaking purely as an outsider—I can’t help wondering whether some of them will eventually hit a wall. So I’ve been watching their progress with interest.
At least for now, those libraries also seem to lag behind mine in rendering quality. I’ve seen cases where layouts break quite dramatically, as well as files that simply cannot be displayed at all. To be frank, that makes me wonder whether the range of practical applications for them is still fairly limited.
As I mentioned earlier, I think the key in this field is figuring out how to efficiently collect as wide a variety of sample documents as possible—and, even better, documents that actually break your renderer. To do that, perhaps you first need to secure a sufficiently large user base.
Where any of these projects, including my own, ultimately ends up is impossible to predict given the current pace of AI progress. Still, at a time when the idea that generative AI will “burn software engineering jobs to the ground” is increasingly becoming conventional wisdom, I suspect the number of software engineering jobs will indeed be reduced substantially. But I also find myself wondering whether somewhere in these differences between projects lies a clue to the kind of value that will remain for humans.
This update has gotten quite long, but one final topic: GPT-6 Astra was released on September 3.
As usual, I follow a policy of using the highest-end model available for difficult tasks, mostly for my own peace of mind, so I’ve started assigning the harder work to Astra. That said, I already had essentially no complaints with Sol. For this particular task, model capability may have effectively hit the ceiling, to the point where I’m not sure there is much practical benefit to using a higher-end model anymore. Still, if it’s available, I might as well use it.
So far, at least for this project, Astra mostly seems to consume more tokens. I haven’t yet had a moment where it did something that made me think, “Wow, that was remarkable.”
One last note before ending this update.
The title of this article says I built this with Claude, but lately I’ve started to feel slightly self-conscious about the fact that, in practice, it has now become a GPT project.
For a while, the prevailing mood seemed to be something like, “Anthropic is amazing; OpenAI is finished.” Now OpenAI has thoroughly overtaken Anthropic in both model capability and cost-performance—and that seems to have become the broader consensus as well. It really drives home just how impossible this world is to predict.
Oh, and by the way, this entire article was written by a human. That’s my policy, and I intend to keep it that way.