Skip to content

add production monitoring best practices - #3745

Open
id wants to merge 6 commits into
release-6.3from
20260806-add-monitoring-best-practices
Open

add production monitoring best practices#3745
id wants to merge 6 commits into
release-6.3from
20260806-add-monitoring-best-practices

Conversation

@id

@id id commented Aug 6, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

@id id added this to the 6.3.0 milestone Aug 6, 2026
Comment thread en_US/observability/log.md Outdated
because fields can contain client IDs, usernames, topics, peer addresses, and
error details.

Monitor the collection path itself. Alert if a node stops sending logs while it

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no log could mean the system is idle, or no error/warning

Comment thread dir.yaml Outdated
title_cn: Production Monitoring Best Practices
title_ja: Production Monitoring Best Practices
path: observability/monitoring-best-practices
lang: en

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is new ?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is new doc, yes

zmstone
zmstone previously approved these changes Aug 6, 2026
thalesmg
thalesmg previously approved these changes Aug 6, 2026
@Meggielqk
Meggielqk dismissed stale reviews from thalesmg and zmstone via db6a21f August 7, 2026 04:56

- The collector or transport is unhealthy.
- The collector or transport rejects or drops records.
- The central backend approaches its storage or retention limits.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@id What does “retention limits” specifically refer to here?What operational condition should trigger the alert?

  • The backend is approaching its storage capacity or a retained-data quota; or
  • Storage pressure means the backend might no longer satisfy the required log retention period.

@Meggielqk

Copy link
Copy Markdown
Collaborator

@id @zmstone, I reorganized and refined the content. Please do the final proofreading for the EN and ZH docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants