Skip to content

DeadSystemException on main thread while BNatives native check is running (18.3.0) #80

Description

@eby-dev

Describe the bug
Since integrating freeRASP (TalsecSecurity-Community:18.3.0) into our production app, we see fatal android.os.DeadSystemException crashes on the main thread at random framework call sites. Before freeRASP was added, our in-app crash reports contained no DeadSystemException at all; they started appearing right after the release that integrated it.

In every crash thread dump we checked, a freeRASP thread was actively running native code at the moment of the crash, while most other threads were idle.

Main-thread crash sites (all Caused by: android.os.DeadSystemException):

  • ActivityThread.handleUnbindService / handleStopService / handleCreateService (e.g. Firebase JobInfoSchedulerService, TagManagerService)
  • ActivityThread.handleSleeping
  • InputMethodManager.removeImeSurface
  • ActivityClient.activityStopped, TextServicesManager.finishSpellCheckerService, AudioManager.playSoundEffect

To Reproduce
Not reproducible on our own test devices. One affected user (POCO 21121210G, Android 12) hits it consistently: the app crashes almost every time the soft keyboard is dismissed (e.g. after typing in a search field). Other users hit it sporadically.

Expected behavior
freeRASP checks should not cause, or coincide with, DeadSystemException crashes on the app's main thread.

Screenshots
freeRASP thread at the moment of crash, seen in 3 of 5 DeadSystemException dumps (Android 9, 12, 13):

DefaultDispatcher-worker-N:
  at androidx.security.BNatives.s(SourceFile)
  at com.aheaditec.talsec.security.e2.m(e2.java:2)
  at com.aheaditec.talsec.security.e2.n(e2.java:1)
  at com.aheaditec.talsec.security.f.a(SourceFile:34)
  at com.aheaditec.talsec.security.e2.k(e2.java:1)
  at com.aheaditec.talsec.security.k4$c.invokeSuspend(SourceFile:6)

Periodic check (Android 10, crash in handleSleeping):

pool-23-thread-1:
  at androidx.security.BNatives.r(SourceFile)
  at com.aheaditec.talsec.security.e2.l(e2.java:2)
  at com.aheaditec.talsec.security.e2.i(e2.java:3)
  at com.aheaditec.talsec.security.f.a(SourceFile:34)
  at com.aheaditec.talsec.security.e2.g(e2.java:3)
  at com.aheaditec.talsec.security.i4.run(i4.java:6)
  at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:307)

Initialization (Android 12, crash in handleCreateService):

DefaultDispatcher-worker-7:
  at java.lang.Class.getName(Class.java:1048)
  at com.aheaditec.talsec.security.z3.a(SourceFile:18)
  at com.aheaditec.talsec.security.z3.<init>(SourceFile:1)
  at com.aheaditec.talsec.security.d2.<init>(SourceFile:1)
  at com.aheaditec.talsec.security.l2.<init>(SourceFile:1)
  at com.aheaditec.talsec.security.n4.<init>(n4.java:1)
  at com.aheaditec.talsec.security.j4.a(SourceFile:19)

For comparison, in unrelated crashes (ClassCastException, IllegalStateException) the active freeRASP thread was in ApkVerifier.verify or PackageManager.getPackageInfo; BNatives.s did not appear there.

Please complete the following information:

  • Device: POCO 21121210G, Redmi 23108RN04Y, multiple Samsung devices (~91% of one crash group)
  • OS version: Android 9 (most affected), 10, 12, 13
  • Version of freeRASP: 18.3.0 (Community)

Additional context

  • Started via Talsec.start(context, config) (default mode) from Application.onCreate(), with prod(true). Kotlin 2.0.21, AGP 8.7.3, compileSdk/targetSdk 36, minSdk 24, R8 enabled in release.
  • system_server appears to still be alive when this happens: right after the crash, our uncaught-exception handler successfully starts a new activity in a fresh process. So we suspect the DeadObjectException comes from the binder failing inside our own process, not from a real system restart.
  • We know 18.0.4 patched getInstalledPackages throwing DeadSystemException.

Questions:

  1. Is this a known issue, or related to what BNatives.s / BNatives.r do (e.g. heavy binder/ioctl usage exhausting the process binder buffer)?
  2. Do the 19.2.x fixes ("ANR during SDK initialization caused by blocking I/O", "background SDK initialization stalling", "native crash caused by std::terminate() race condition") address this?
  3. Would starting with TalsecMode.BACKGROUND reduce the risk?

Full thread dumps are available on request.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions