Loading video...

Video Failed to Load

Go Home

#SakuraWars2 DevLog: Dealing with Images This game contains a fair bit of text within images. So for a translation patch, the images need to be extracted, redone, and patched.

19,529 views • 1 year ago •via X (Twitter)

10 Comments

NSteam 🇵🇸's profile picture
NSteam 🇵🇸1 year ago

This includes font sheets, all battle UI, chapter intro screens, battle screens, boss banners, minigame images, theater room names, system menus, menu backgrounds, loading screens, splash screen, game title screen, and other misc images. Currently amounting to 1,735 images!

NSteam 🇵🇸's profile picture
NSteam 🇵🇸1 year ago

For some of these, like the room names in the theater & names appearing within battle UI, there were hundreds of images. To do them by hand was tedious so I wrote a tool to auto-generate them given the translated names in a text file. Some of the battle UI was done by hand though

NSteam 🇵🇸's profile picture
NSteam 🇵🇸1 year ago

We got the help of some very talented artists to help redo some of the more complicated images like these boss images, chapter screens, and minigame images.

NSteam 🇵🇸's profile picture
NSteam 🇵🇸1 year ago

Most images are stored as raw vdp1 format image data. Others are tiled images displayed in vdp2. These are easy to extract using tools I wrote. Others are compressed using PRS compression which was a standard across many Sega titles. These are a bit more complicated.

NSteam 🇵🇸's profile picture
NSteam 🇵🇸1 year ago

Some however are encoded using some bizarre encoding algorithm. It doesn't compress the image, just obfuscates it. The Kinematron, some minigames, and credits all use this. And these are a pain.

NSteam 🇵🇸's profile picture
NSteam 🇵🇸1 year ago

To extract these, the encoded data needs to be found, the encoding algorithm needs to be figured out(which Ghidra helps with) and then it needs to be reverse engineered so that I can re-encode the translated images and this part was pretty tricky.

NSteam 🇵🇸's profile picture
NSteam 🇵🇸1 year ago

Above image is full of artifacts unless it is decoded. Here is one of the images properly decoded. What makes it worse is that there are multiple encoding schemes. The kinematron & minigames use one, credits use another. I call this algorithm the "Kinematron Encoding Algorithm"

NSteam 🇵🇸's profile picture
NSteam 🇵🇸1 year ago

I have yet to reverse engineer the credits algorithm. Not sure if it will make it into the final patch. Maybe! But for now, everything else we have encountered has been extracted, reworked, and patched =)

burntends | Macross Do You Remember Love?'s profile picture
burntends | Macross Do You Remember Love?1 year ago

Always great to see more news

Norbjunior's profile picture
Norbjunior1 year ago

@burntends2 Super cool to learn how some of this works!! Thank you for all the hard work!!!!

Related Videos

If you're building a PDF RAG pipeline: Should you be using OCR and 𝘁𝗲𝘅𝘁-𝗯𝗮𝘀𝗲𝗱 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹 methods, or just 𝗲𝗺𝗯𝗲𝗱 𝗶𝗺𝗮𝗴𝗲𝘀 𝗱𝗶𝗿𝗲𝗰𝘁𝗹𝘆 using late interaction models? This paper says the answer might actually be 𝘣𝘰𝘵𝘩. My colleagues at Weaviate released IRPAPERS, a benchmark comparing 𝗶𝗺𝗮𝗴𝗲-𝗯𝗮𝘀𝗲𝗱 and 𝘁𝗲𝘅𝘁-𝗯𝗮𝘀𝗲𝗱 retrieval over 3,230 pages from 166 scientific papers. The setup: Take the same PDFs and process them two ways. For text, run OCR with GPT-4.1 and embed with Arctic 2.0 + BM25 hybrid search. For images, embed raw page images with ColModernVBERT multi-vector embeddings. Test both on 180 needle-in-the-haystack questions. 𝗧𝗵𝗲 𝗿𝗲𝘀𝘂𝗹𝘁𝘀: Text edges out images at the top rank: 46% vs 43% Recall@1 But images match or exceed text at deeper recall: 93% vs 91% Recall@20 But text and image based methods actually fail on 𝘥𝘪𝘧𝘧𝘦𝘳𝘦𝘯𝘁 𝘲𝘶𝘦𝘳𝘪𝘦𝘴. At Recall@1: • 22 queries succeed with text but fail with images • 18 queries succeed with images but fail with text This complementarity is what makes 𝗠𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹 𝗛𝘆𝗯𝗿𝗶𝗱 𝗦𝗲𝗮𝗿𝗰𝗵 work. By fusing scores from both text and image retrieval, they achieved: • 49% Recall@1 (beating either modality alone) • 81% Recall@5 • 95% Recall@20 More in the video below 🔽 Dataset: Paper: Code:

Victoria Slocum

44,145 views • 5 months ago