# Genkō Counter: self-hosted website, OCR and photo copies

This package includes the website and a Node server that runs OCR through a local Ollama vision model. No paid API or API key is required. Uploaded, pasted and camera photos are saved automatically on your server, independently of OCR. Saved copies are compressed in the browser before upload: JPEG at 85% quality, targeting a 2,200-pixel long edge. If re-encoding an uploaded file would increase its byte size, the smaller original is kept. Counting and OCR use their existing image pipeline. On-device OCR remains available. You can use photo saving and box counting without installing an OCR model.

## Quick start with Docker

Unzip the package, open a terminal in its directory, then run:

```sh
docker compose up -d --build
docker compose exec ollama ollama pull qwen3-vl:8b
```

Open **http://localhost:3000** on the server. The website and photo archive are ready immediately; the Ollama pull command is only needed for server OCR. Choose **Server**, then **Check server** for OCR. The first model download is about 6.1 GB. Opening the page starts loading the model in the background; the first recognition may still take longer while it finishes loading. The model stays in memory for 15 minutes after a scan. Docker uses CPU by default; a supported GPU makes a substantial difference. Handwriting still needs review and local CPU OCR may be slow.

For a smaller model, create a `.env` file containing `OCR_MODEL=qwen3-vl:4b`, pull it with `docker compose exec ollama ollama pull qwen3-vl:4b`, and run `docker compose up -d` again. For an NVIDIA GPU, add `gpus: all` to the `ollama` service after installing the NVIDIA Container Toolkit on your Docker host. Use a current Docker Compose version supporting that setting. On a Mac, native Ollama can use the GPU; use the non-Docker setup below.

Stop with `docker compose down`. Photos are retained in the host `photos` folder, and model downloads are retained in the Docker volume. Rebuilding or replacing the container keeps both. Back up the host photo folder separately.

## Small Ubuntu server: website and saved photos

For a server with about 1 GB RAM, use the supplied `deploy/install-ubuntu.sh` instead of Docker Compose. It installs Node 22 and Caddy, runs the website as an unprivileged system service, keeps the app private on port 3000, enables compressed photo copies, and disables server OCR (`OCR_ENABLED=false`). The website automatically uses on-device OCR. No Docker or vision model is installed.

Point the domain's IPv4 record at the server, clear any nonworking IPv6 record, and allow inbound TCP 80 and 443 in the cloud firewall. On the server, run:

```sh
sudo bash deploy/install-ubuntu.sh genkoucount.duckdns.org 34.40.229.145
```

The installer checks for existing web services before making changes, refuses to replace unrelated Caddy configuration, and verifies the Node binary against official release checksums. Caddy handles HTTPS certificate renewal. The script prints a generated login for user `aaron`, kept in `/root/genkoucount-login.txt`; use it when opening the site. Password protection covers the website and uploads. Photos are private files under `/var/lib/genkou-counter/photos`, separate from the app at `/opt/genkou-counter`. No existing photos are recompressed or deleted.

To update, rerun the installer; it downloads the current self-host package and preserves saved photos and the generated login. Inspect status with `sudo systemctl status genkou-counter caddy` and logs with `sudo journalctl -u genkou-counter -u caddy`. The installer checks the running app locally; it cannot confirm external firewall access or completed public certificate issuance. If HTTPS fails, inspect the Caddy logs and DNS/firewall settings.

## Without Docker

Install Node.js 22 or newer and Ollama 0.12.7 or newer. Start Ollama, then run:

```sh
ollama pull qwen3-vl:8b
node server/index.mjs
```

Open **http://localhost:3000**. No npm install is needed. The default server listens only on your machine. Environment variables: `PORT` (3000), `HOST` (127.0.0.1), `OLLAMA_URL` (http://127.0.0.1:11434), `OCR_MODEL` (qwen3-vl:8b), `OCR_TIMEOUT_SECONDS` (180, maximum 600). Photo variables are documented below.

## Use a phone or the published website

The easiest arrangement is to host this entire website and OCR server together behind an HTTPS reverse proxy. Set `PUBLIC_ORIGIN=https://your-server.example` and pass traffic to port 3000. HTTPS is required for a phone's live camera. Uploading a photo is also available. Keep Ollama's port 11434 private. For access outside your own network, put the entire site and API behind access control at your reverse proxy. Origin checks are not user authentication.

To connect the published Genkō Counter instead, set:

```text
PUBLIC_ORIGIN=https://your-server.example
ALLOWED_ORIGINS=https://genkou-box-counter.aaron-ye1.chatgpt.site
```

Restart the server. On the published site, expand **OCR settings**, enter `https://your-server.example` as the server address, and press **Check server**. The address is saved in that browser. Use an HTTPS origin without a path; the app adds `/api/ocr` and `/api/ocr/status`. A blank address uses the server serving the website. Localhost HTTP is allowed, but browsers may ask permission for access to a local server. If a browser blocks it, open the locally served website directly.

Docker's default port binding is localhost only. For a reverse proxy running elsewhere, set `BIND_ADDRESS=0.0.0.0` in `.env` and restrict port 3000 to your LAN/proxy. Set `PUBLIC_ORIGIN` to the exact URL you use, including its port for LAN HTTP, then restart with `docker compose up -d`. Live camera access still requires HTTPS on a phone; file uploads work over LAN HTTP. `PORT` changes the published host port. `ALLOWED_ORIGINS` accepts a comma-separated list of exact origins, never a wildcard. The API requires an allowed Origin on POST, accepts image data only (no remote image URLs), limits OCR JSON requests to 6 MiB and original photo files to 32 MiB by default, and runs one scan at a time. This is a personal server; add authentication at the proxy for shared/public hosting.

## Saved photo copies

Saving does not wait for OCR, and a failed grid scan does not discard the photo. Saving runs silently in the background, with one attempt per photo and no on-page notices or retry controls. A failed attempt is not retried or retained in the browser. The public hosted website does not have an archive itself; it can save to your configured self-hosted server after that server has been checked and allows the hosted website's origin.

Docker defaults to `./photos` beside `compose.yaml`, mounted as `/data/photos` inside the container. Native Node defaults to `data/photos` in the project directory. Copies are organized as:

```text
photos/YYYY-MM-DD/<unique-id>/photo.jpg
photos/YYYY-MM-DD/<unique-id>/info.json
```

The extension follows the stored image type: compressed copies use `.jpg`; uncompressed fallbacks keep their original type. `info.json` records the original filename, upload/camera/paste source, UTC timestamp, stored byte length, SHA-256 hash, whether compression was applied, and the original upload byte length (when available). Filenames never select a server path. Photos and metadata are stored outside the public website folder; they have no public download URL. Use the server filesystem to retrieve or delete them. There is no automatic deletion. Dates use UTC. A browser reload can interrupt a pending upload. Check the server photo folder to confirm copies; no success or failure banner appears on the page.

Supported archive formats: JPEG, PNG, WebP, GIF, BMP, AVIF and HEIF/HEIC. The browser must also be able to open the image for counting. SVG files are not archived. Compression uses the whole photo, not a grid crop or OCR-cleaned image. It does not alter or delete the original file on your device. The archive stores one copy per capture/upload, not both original and compressed versions. Compression affects new uploads only; existing saved photos are unchanged.

Optional `.env` settings for Docker:

```dotenv
# Persistent folder on the server; an absolute path is also supported.
PHOTO_PATH=./photos
SAVE_PHOTOS=true
PHOTO_MAX_MB=32
PHOTO_COMPRESS=true
PHOTO_MAX_EDGE=2200
PHOTO_JPEG_QUALITY=85
# Set this when using an HTTPS reverse proxy.
# PUBLIC_ORIGIN=https://counter.your-domain.example
```

Compression is enabled by default. Set `PHOTO_COMPRESS=false` to keep uploaded originals byte-for-byte and camera captures at 95% JPEG quality. `PHOTO_MAX_EDGE` accepts 800–5,000 pixels and `PHOTO_JPEG_QUALITY` accepts 60–95. Settings take effect after restarting the server and reloading the website. Unsupported browser decoding falls back to the original upload. There are no compression controls or notices in the website.

For native Node, use `PHOTO_DIR=/absolute/path/to/photos` instead of `PHOTO_PATH`. `PHOTO_MAX_MB` supports 1–100 MB. `SAVE_PHOTOS=false` disables saving. Ensure the process can write its photo directory and has enough disk space. Docker's default process may create root-owned files; use administrator access to back up or manage them if needed. To move an existing archive, stop the service, copy its folder, change the path, and restart.

## Update an existing installation

Back up the photo folder. Replace the application files with this download, keeping your `.env` and photo folder, then run `docker compose up -d --build`. Native Node users should restart `node server/index.mjs`. Replace both the website and server together so the photo upload endpoint and browser code match. Do not copy the archive into `dist` or another public web directory.

## Troubleshooting

- **Model not ready:** run the `ollama pull` command and check `docker compose logs ollama` or the native Ollama process.
- **Not enough memory / too slow:** use the 4b or 2b vision model, a smaller section, or a supported GPU. Accuracy and speed depend on hardware and handwriting.
- **Camera unavailable:** use HTTPS on a phone, or choose Upload.
- **Photo not saved:** check the photo directory permissions, disk space and `PHOTO_MAX_MB`. The browser attempts each photo once; a new upload is needed after correcting a server problem.
- **Connection blocked:** check HTTPS, `PUBLIC_ORIGIN`, `ALLOWED_ORIGINS`, and your reverse proxy. Server addresses cannot include `/api/ocr`.

The server/API plumbing and photo saving were tested through real HTTP requests and filesystem writes, with a mocked Ollama model. Actual model inference, hardware performance and handwriting accuracy were not tested on this machine.

Official references: https://ollama.com/library/qwen3-vl:8b and https://docs.ollama.com/api/chat.

## Freeform pages and count filters

Choose **Auto-detect** or **Freeform text** for plain or lined paper. Choose the OCR language (Japanese, simplified/traditional Chinese or English). Freeform counts use the editable recognition result; check it and remove unwanted notes or crossed-out text. You can also paste text without a photo.

Freeform toggles recount edited text immediately. Grid toggles filter individual boxes and never replace the optical count with OCR text. Compact punctuation is classified optically; other types use confident, geometrically aligned on-device OCR symbols or the type chosen in the box inspector. Server text is not assigned to boxes without alignment. Unclassified boxes stay counted. Manual Count corrections and manually entered addition counts remain as entered. Outside-grid optical suggestions need review and inclusion; tap their preview to add/remove character rectangles. They are estimates, not guaranteed character recognition. Spaces and line breaks never count.

If updating an older self-hosted server, replace both the website and server with this download to support freeform/language requests.
