↓ Skip to main content
  1. Home projects/

comments-tool: finding Japanese comments in C/C++

·

My own tool, a bit silly, but it does the job. Some code bases come with comments in Japanese. If you do not read Japanese, you first need to know where they are and what they say. comments-tool walks through a directory, finds the Japanese comments in C and C++ files and puts them into one JSON file, so they can be translated in one go.

The code is on GitHub.

What it does
#

It reads .c, .cpp, .h and .hpp files, in a directory or one file. It finds // and /* ... */ comments and keeps those with hiragana, katakana or kanji. Each one goes to translations.json as a key with an empty value. You fill in the translation later, by hand or with any translator. A comment already in the file is not added twice.

It can also list every Japanese character with its line and position, list non-ASCII characters, and find and remove the BOM character.

It does not translate anything and does not change your code, except for --remove-bom. The regular expressions are simple, so a // inside a string literal counts as a comment too.

How to run it
#

You need Python 3.8 or newer. With uv:

git clone https://github.com/marcinklimek/comments-tool.git
cd comments-tool
uv venv && source .venv/bin/activate
uv pip install -e .

comment_tool --directory path/to/source   # collect Japanese comments
comment_tool --scan --file path/to/file.cpp   # show where the Japanese characters are
comment_tool --scan-bom                   # find BOM characters
comment_tool --remove-bom                 # remove them

Without a flag the tool collects comments. It writes translations.json and translation.log in the current directory.

Example
#

For a file with this line:

int speed = 10; // 速度の初期値

the run adds this to translations.json:

{
  "// 速度の初期値": ""
}

The log says in which file and on which line the comment was found. The empty value is where the translation goes (“initial value of the speed”).

Related