Skip to content

fix: user data corruption on unclean Windows shutdown - #213

Merged
edospadoni merged 1 commit into
mainfrom
fix/user-data-corruption
Sep 18, 2026
Merged

edospadoni merged 1 commit into
mainfrom
fix/user-data-corruption

Conversation

@edospadoni

@edospadoni edospadoni commented Sep 16, 2026 •

Copy link
Copy Markdown
Member

The problem

After an unclean Windows shutdown, user_data.json and available_user_data.json come back at roughly the right size but filled entirely with 0x00 → NethLink hangs forever on "Starting application...".

Today the only way out is renaming both files by hand and logging in again.

Two causes

1. Non-atomic writes — saveToDisk() wrote straight to the destination with writeFileSync, never flushing. NTFS journals metadata only: on power loss the new file size is replayed from the journal while the data pages, still in the page cache, are lost → a file of the new length, all zeros.

2. A read that throws — getAvailableFromDisk() called JSON.parse with no try/catch. The SyntaxError escaped startApp(), which nobody awaits and which had no handler → showLogin() and the splash screen teardown were never reached.

Confirmed in the field: in the first test, renaming only available_user_data.json was enough. That matches — getFromDisk() (which reads user_data.json) was already wrapped in try/catch, getAvailableFromDisk() was not.

The fix

Writes .tmp → fsync → rename onto the target, keeping a .bak of the previous version. The fsync before the rename is the load-bearing part: without it, a NUL-filled file can still show up under the final name.
Reads Reject empty, NUL-filled and unparsable files; fall back to the .bak; never throw.
Startup startApp() is guarded: on error it degrades to the login window instead of freezing. unhandledRejection and uncaughtException are logged.

The .bak is consulted only when the primary file exists but is unusable, so deleting the json files by hand still forces a clean login, exactly as today.

Already broken machines recover on their own after updating: they reach the login screen instead of a frozen splash, with no manual file renaming. The data format is unchanged, so rolling back is safe.

How to test

  1. Set up NethLink and leave it running in the system tray.
  2. Power off the machine with the power button.
  3. On reboot, NethLink must start normally.

To exercise the second layer without a power cut: fill both json files in %APPDATA%\nethlink\ with zeros and restart NethLink → it must reach the login screen (not a frozen splash), and with the .bak files present it must come straight back up without asking for credentials.

@edospadoni edospadoni self-assigned this Sep 16, 2026
NethLink could hang forever on the "starting application" splash screen
after a power loss, with both user_data.json and available_user_data.json
recovered at roughly the right size but filled with 0x00.

Two independent causes:

- saveToDisk() wrote straight to the destination with writeFileSync and
  never flushed. NTFS journals metadata only, so after a power loss the
  new file size is replayed from the journal while the data pages are
  still in the page cache and lost, leaving a NUL filled file. Writes now
  go to a temp file, get fsync'd and are renamed over the target, so the
  previous content stays intact until the new one is durable on disk, and
  a .bak of it is kept.

- getAvailableFromDisk() parsed the file with no try/catch, so the
  SyntaxError propagated out of the fire and forget startApp(), where
  nothing caught it: showLogin() and the splash screen teardown were
  never reached. Reads now reject empty, NUL filled and unparsable files,
  fall back to the backup and never throw, startApp() is guarded, and
  unhandled rejections are logged.

Reads consult the backup only when the primary file exists but is
unusable, so removing the json files by hand still forces a clean login.

Already affected installations recover on their own: they land on the
login screen instead of a frozen splash, with no manual file renaming.
@edospadoni
edospadoni force-pushed the fix/user-data-corruption branch from 73b48f0 to 1b95956 Compare September 16, 2026 07:38
@github-actions

Copy link
Copy Markdown

Automatic builds from https://github.com/NethServer/nethlink/actions/runs/35069584704.
Commit: 1b95956

Name Platform Link
win-app.exe Windows (x64) Link
macos-app-x64.dmg MacOS (x64) Link
macos-app-arm64.dmg MacOS (arm64) Link
linux-app.AppImage Linux (x64) Link

@edospadoni
edospadoni merged commit 7c50643 into main Sep 18, 2026
4 checks passed
edospadoni added a commit that referenced this pull request Sep 18, 2026
ESLint was entirely broken on main: .eslintrc.cjs extended 'react-app'
and 'react-app/jest' while eslint-config-react-app was never installed,
so 'npm run lint' failed before linting a single file.

Rather than repair the eslintrc setup on eslint 8, this migrates to
eslint 9 flat config, which also unblocks the two Renovate majors that
were stuck on the 'eslint >= 9' peer requirement:

- eslint ^8.56.0 -> ^9.39.5
- @electron-toolkit/eslint-config-ts ^1.0.1 -> ^3.1.0 (#206)
- @electron-toolkit/eslint-config-prettier ^2.0.0 -> ^3.0.0 (#205)
- eslint-plugin-react ^7.33.2 -> ^7.37.5 (flat config support)
- eslint-plugin-react-hooks ^5.2.0 (was referenced, never installed)

.eslintrc.cjs and .eslintignore are replaced by eslint.config.mjs.

Config decisions:

- React and react-hooks rules are scoped to src/renderer/**. The main
  process has plain functions named use* (useNethVoiceAPI, useLogin)
  that rules-of-hooks reported as hook violations in AccountController.
- camelcase is off. Every one of its 41 reports was a snake_case field
  fixed by the backend API contract (speeddial_num, shared_groups,
  dst_cnam) or a theme token (extra_large, full_w).
- react/prop-types is off, since TypeScript already checks props.
- spaced-comment carries markers: ['/'] so --fix cannot mangle
  TypeScript triple-slash directives into '// / <reference ... />',
  which silently broke the vite/client types behind every .svg import.

Beyond formatting, these lint errors needed real fixes:

- useNethVoiceAPI.ts: drop the redundant async Promise executor around
  logout; the body already resolved unconditionally in finally, so
  behaviour is unchanged and no rejection path is lost.
- logger.ts: replace the unsafe Function type with an explicit signature.
- main.ts, Modal/index.tsx: merge duplicate imports.
- AboutModule.tsx: drop the empty destructuring pattern.
- usePresenceService.ts: drop an empty finally block.
- ipcEvents.ts: document the intentionally silent catch in the drag path.
- store.ts: move a // @ts-ignore next to the expression it suppresses.
  Reformatting split the statement it used to cover onto two lines,
  which silently un-suppressed a zustand typing error.

'npm run lint' now reports 0 errors and 339 warnings. typecheck and
build are green, and the code merged from #192 and #213 is intact.

Supersedes #215.
@edospadoni
edospadoni deleted the fix/user-data-corruption branch September 25, 2026 07:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant