Reading dmesg on Embedded Without an RCU Deep Dive

On embedded Linux, the kernel ring buffer often fails you in three boring ways: too much noise, unclear when a line landed, and no sense of whether the same failure is looping. You do not need an RCU or scheduler deep dive for a first pass. Log levels · timestamps · repeat patterns are usually enough.

This post stays on three axes only: how to read levels · how to read timestamps · how to treat repeating panics. No invented board diaries, prices, or affiliates.

Grounded in dmesg(1) (util-linux) and the kernel’s Message logging with printk.

How do you read levels?

One-line answer: printk priorities run 0 (emerg) … 7 (debug); lower numbers are more severe. Whether a line hits the console depends on comparing that priority to console_loglevel. When reading, narrow with dmesg -l / -x / -n.

printk level table (summary):

NameStringAlias
KERN_EMERG"0"pr_emerg()
KERN_ALERT"1"pr_alert()
KERN_CRIT"2"pr_crit()
KERN_ERR"3"pr_err()
KERN_WARNING"4"pr_warn()
KERN_NOTICE"5"pr_notice()
KERN_INFO"6"pr_info()
KERN_DEBUG"7"pr_debug(), etc.

Practical reading patterns from dmesg(1):

# Errors and warnings only (example)
dmesg --level=err,warn

# err and more severe (crit/alert/emerg) — util-linux "+" syntax
dmesg --level=err+

# Decode facility/priority numbers to human prefixes
dmesg -x

# Change console print threshold only (does not print/clear the buffer)
dmesg -n 5
# or: write to /proc/sys/kernel/printk  (printk docs)

Caveats:

  1. Filter (-l) and console level (-n) are different axes. -l controls what you display now; -n controls what may hit the console later.
  2. Landing in the ring buffer is separate from immediate console output. printk docs: a message prints to the console when its priority is higher than (numerically lower than) console_loglevel.
  3. BusyBox dmesg may expose fewer flags than util-linux. Check dmesg --help on the target.
  4. Permission denied often comes from dmesg_restrict (dmesg(1) EXIT STATUS; syslog(2)).

Do not: invent a board rule like “on our SoC, only level 3 matters.” Defaults follow kernel config and the runtime printk quadruple.

What about timestamps?

One-line answer: Default output is usually seconds since boot (raw); human forms use -T/--ctime, ISO uses --time-format=iso, gaps use -d/delta. The man page warns that -T and iso can be wrong after suspend/resume.

dmesg(1) format axes:

OptionMeaningCaution
(default / raw)Seconds since bootCommon for scripts and ordering
-T / --ctimeHuman-readable wall timeMay be inaccurate after SUSPEND/RESUME
--time-format=isoISO-8601-styleSame suspend caveat as ctime
-d / deltaInter-message gapsWith --notime, delta only
-e / --reltimeLocal time + deltaConversion may be inaccurate (see -T)
-t / --notimeHide timestampsBody-only reading

A practical embedded order:

# 1) Skim with boot-relative seconds
dmesg

# 2) Cut to severe lines, show gaps
dmesg --level=err+ -d

# 3) Align with host clocks when comparing (if supported)
dmesg --time-format=iso

# 4) Live follow (kernel ≥3.5, readable /dev/kmsg)
dmesg -w

Reading tips:

  1. For early driver failures, use raw seconds first: which second-range clusters the errors?
  2. Reach for -T/iso only when you need calendar time; on suspend-capable devices, treat man-page warnings as part of the contract.
  3. Ask whether the same symptom arrives in tight deltas or after a long idle gap—that feeds the next section.

How do you handle repeating panics?

One-line answer: Do not stop at one Kernel panic line. Check whether the same call site and preceding err lines repeat on a measurable interval. Narrow with err+, keywords, and -w, and write the repro unit (once per boot / every N seconds / after a given ioctl) before any RCU theory.

Observation checklist (docs and practice; not a SoC anecdote):

[ ] dmesg --level=err+ -d          # severe lines + gaps
[ ] dmesg -x | grep -E 'panic|Oops|BUG:|Internal error'
[ ] 20–50 lines before the first panic/Oops: prior err/warn?
[ ] raw seconds or delta: how often does the same pattern recur?
[ ] dmesg -w during repro if available
[ ] ring buffer overrun? (check buffer size / log collection path)

How to bucket repeats:

PatternWhat you seeNext question
Once per bootSame driver/device strings in a similar second rangeProbe / clock / power / DT?
Burst with short deltasSame err chain, then panic/OopsRetry loop or hard fault?
Long gap, then againAfter idle or a specific workloadWorkload / thermal / bus timeout?
Console quiet, buffer noisyPresent in ring, filtered from consoleCheck -n / /proc/sys/kernel/printk

Be honest about limits:

  1. dmesg is not full root-cause. Stacks, registers, and lockup debuggers are other tools (kgdb, crash dumps, …)—out of scope here.
  2. If an RCU stall line appears, record that a stall occurred; do not chase RCU internals in this pass (brief: no deep dive).
  3. On permission errors, check dmesg_restrict and the calling user first.

One-line wrap-up: Cut by level (-l/-x) → group by time/gap (raw/-d/-T) → write only the repeating panic/Oops unit. First triage does not require an RCU paper.

FAQ

Do BusyBox and util-linux dmesg share the same flags?
Often no. --level=err+, --time-format=iso, and -w are documented for util-linux. Trust the target binary’s --help.

Does raising -n reprint old logs?
No. -n/--console-level only changes the console threshold; it does not print or clear the ring buffer (dmesg(1)).

If -T looks wrong, what should I trust?
Right after boot without suspend, -T/iso is convenient. When in doubt, lock ordering with raw boot seconds and delta. The man page warns about SUSPEND/RESUME inaccuracy.

Is saving one panic line enough?
Usually not. Without preceding err/warn, repeat interval, and repro conditions, the next debug loop stretches. Keep at least an err+ window plus timestamps.

Sources