Auto Remediation Configuration Menu

The submenu is displayed exactly as follows:

The submenu is displayed exactly as follows:

  • Reload Configuration
  • Use SMTP from AppViewX Product Settings
  • View Configuration
  • Prepare / Edit auto_remediation_config.yaml
  • Apply Configuration
  • Run Alert Cycle
  • Generate Report
  • Exit
Option What it does Underlying command
1 Optionally edit auto_remediation.conf, then reload it into operator, cluster-manager, and node-manager. Skipped when unchanged. ./auto_remediation.sh --reload-config
2 Preview SMTP from AppViewX product settings, confirm, and sync to monitoring. ./auto_remediation.sh --sync-smtp --non-interactive
3 Show masked auto_remediation.conf, then export applied Node Manager / Cluster Manager configuration. ./auto_remediation.sh --view-config <type>
4 Create the YAML from the template if missing, optionally open an editor, then validate YAML. Local file edit + YAML validation
5 Apply auto_remediation_config.yaml to Node Manager, Cluster Manager, or both. Skipped when unchanged. ./auto_remediation.sh --apply-config <type> --yes --non-interactive
6 Process pending alert records immediately. ./auto_remediation.sh --run-alert-cycle
7 Generate an HTML or CSV report from RQLite. ./auto_remediation.sh --report --format html|csv
8 Leave the submenu and return to the Interactive IU main menu. No script invocation
Note: Commands are executed from the installer scripts directory. The Interactive IU is the supported customer interface. Commands are listed so Support Engineers can correlate IU actions with installer logs.

Reload Configuration

Purpose

Reload the latest scripts/auto_remediation.conf values into the Auto Remediation operator, Cluster Manager, and Node Manager without reinstalling the platform. You can review and update conf flags first.

When to use it

  • After changing SMTP, alert schedule, cooldown, webhook, watchdog, report, or node-placement values in auto_remediation.conf.
  • After enabling ALERT_MANUAL_TRIGGER_ENABLED if it was set to false.
  • When Support needs to republish conf to running components after a validated conf edit.
Prerequisites
  • Namespace auto-remediation exists and the operator is installed and ready.
  • scripts/auto_remediation.conf exists and is parseable.
  • auto_remediation.conf contains ALERT_CRON_SCHEDULE and RQLITE_DB_RETENTION_DAYS (IU light validation).
  • ConfigMap auto-remediation-operator-config exists (install completed / configuration applied at least once).
  • If Node Manager is installed, sudo on worker nodes is required for Node Manager reload.

Expected output

Then use menu option 3 (View Configuration) and confirm the published values. Passwords remain masked.

Situation Expected result
Conf edit with no flag differences No configuration changes detected in auto_remediation.conf.
Reload with no live differences [INFO] No configuration changes detected. Reload is not required.
Reload with differences Components are updated. [SUCCESS] Configuration reloaded successfully. Backup path is logged.
RQLite node list changed Reload is aborted. RQLite membership cannot be changed by reload; reinstall or redeploy is required.

Troubleshooting

Issue Resolution
Prerequisites fail (operator not ready, conf missing keys) Install or repair Auto Remediation. Restore auto_remediation.conf from auto_remediation.conf-template.template.
Validation failed - existing configuration kept Fill required parameters in auto_remediation.conf, then retry.
Reload aborted because AUTO_REMEDIATION_RQLITE_NODES changed Do not use reload for RQLite topology. Plan a reinstall or redeploy of Auto Remediation.
Component reload failed / rollback ran Read logs/auto_remediation_*.log and the backup directory. Verify pods, then retry after fixing the root cause.

Use SMTP from AppViewX Product Settings

Purpose

Copy SMTP settings from AppViewX product settings into Auto Remediation monitoring, then apply them to the operator so alert emails can be sent. This is the supported IU path. Do not use older standalone SMTP-only procedures as the primary workflow.

How SMTP settings are obtained
  • The IU resolves effective SMTP through scripts/script_util/smtp_config.py (resolve-effective).
  • Product settings are read from the AppViewX application database (appSettings properties).
  • If SMTP_CONFIG_SOURCE=manual and manual SMTP is already complete, product-settings sync does not overwrite it.
  • If manual SMTP is incomplete or the source is not manual, product SMTP is written into auto_remediation.conf and SMTP_CONFIG_SOURCE is set to app_db.
  • The operator is updated (smtp-credentials secret and Helm SMTP/alerting values).
Which values are synchronized
Field Synchronized Displayed in IU preview
SMTP_SERVER Yes Yes
SMTP_PORT Yes Yes
SMTP_USER Yes Yes
SMTP_FROM Yes Yes
SMTP_TO Yes Yes
SMTP_PASSWORD Yes (vault-encrypted in conf; Kubernetes secret for the operator) Masked as ******** or (not set)
SMTP_CONFIG_SOURCE Set to app_db when product settings are applied Yes
Prerequisites
  • Namespace auto-remediation exists. Deployment auto-remediation-operator exists.
  • scripts/auto_remediation.conf exists.
  • SMTP is configured in the AppViewX web UI: Settings → General → SMTP (server, from address, recipients, and password as required by that UI). Verify the exact menu labels in the installed AppViewX version if they differ.
Procedure
  1. Select 2. Use SMTP from AppViewX Product Settings.
  2. If SMTP_CONFIG_SOURCE=manual, read the note that product sync will not overwrite complete manual SMTP.
  3. Review the preview of product SMTP (password masked).
  4. Confirm Apply these SMTP settings to Auto Remediation? (Yes/no):. The default is Yes.
  5. The IU runs auto_remediation.sh --sync-smtp --non-interactive.
  6. Confirm the redacted auto_remediation.conf SMTP summary printed after sync.
Failure scenarios
Failure Typical cause What to do
SMTP from AppViewX Product Settings is missing or incomplete Settings -> General -> SMTP is not fully configured. Configure SMTP in the AppViewX UI, then retry option 2.
Operator deployment not found Auto Remediation not installed. Install Auto Remediation, then retry.
SMTP_CONFIG_SOURCE=manual and SMTP is complete Product sync is intentionally skipped for overwrite. Use Reload Configuration to edit monitoring SMTP, or complete or clear manual SMTP if product SMTP should be used.
Sync finished but SMTP is still incomplete Password or required fields still missing after sync. Fix product SMTP. Retry option 2 before Run Alert Cycle.
Command non-zero exit Helm or secret apply failed, or application database unreachable. Check logs/auto_remediation_*.log. Verify application database connectivity and operator health.
Note: The installer also registers an hourly host cron that runs SMTP sync when monitoring-specific manual SMTP is not set. IU option 2 is the on-demand path after you change product SMTP.

View Configuration

Purpose

Display the current Auto Remediation configuration: a masked summary of auto_remediation.conf, then the currently applied Node Manager and/or Cluster Manager configuration collected from the environment.

Configuration file location
Item Location
Runtime flags scripts/auto_remediation.conf
Check YAML scripts/auto_remediation_config.yaml
Exported view output logs/auto_remediation_config_<timestamp>/
Procedure
  1. Select 3. View Configuration.
  2. Review the masked auto_remediation.conf sections.
  3. At View Applied Component Configuration, choose:
  4. The IU runs auto_remediation.sh --view-config <node-manager|cluster-manager|all>.
Note: 1. Node Manager 2. Cluster Manager 3. All Components 4. Back
Configuration sections displayed
Section Keys shown (when present)
SMTP SMTP_CONFIG_SOURCE, SMTP_SERVER, SMTP_PORT, SMTP_USER, SMTP_FROM, SMTP_TO, SMTP_ERROR_TO
Alerting GROUP_BY, ALERT_MIN_COUNT, ALERT_EMAIL_SUBJECT, ALERT_CRON_SCHEDULE, ALERT_MANUAL_TRIGGER_ENABLED, COOLDOWN_PERIOD, COOLDOWN_OVERRIDES, AUTOFIX_ESCALATION_COUNT, AUTOFIX_ESCALATION_WINDOW_MINUTES, WEBHOOK_ENABLED, WEBHOOK_URL
Monitoring / Nodes MONITORING_INTERVAL, AUTO_REMEDIATION_MONITORING_NODES, AUTO_REMEDIATION_RQLITE_NODES, RQLITE_DB_RETENTION_DAYS, CLUSTER_WATCHDOG_ENABLED, CLUSTER_WATCHDOG_CRON_SCHEDULE
Reports REPORT_CRON_SCHEDULE, REPORT_OUTPUT_DIR, REPORT_FORMAT
Sensitive SMTP_PASSWORD_ENCRYPTED shown as (set, hidden) or (not set)

Expected output

Masked conf summary, then collected YAML or config files under logs/auto_remediation_config_<timestamp>/. The IU then prints View configuration completed.

Validation
  • The export directory is created and contains node_manager_config_<ip>.yaml and/or cluster_manager_config.yaml depending on the type selected.
  • Command exit code 0. Failures are logged in logs/auto_remediation_*.log.

Prepare/Edit auto_remediation_config.yaml

Purpose

Create scripts/auto_remediation_config.yaml from the shipped template if it does not exist, optionally edit it in a terminal editor, and validate YAML before Apply Configuration.

Configuration file location

File Path
Working configuration scripts/auto_remediation_config.yaml
Template scripts/auto_remediation_config-template.yaml.template
How to prepare and edit the configuration
  1. Select 4. Prepare / Edit auto_remediation_config.yaml.
  2. If auto_remediation_config.yaml is missing, the IU copies the template and replaces <ENABLE_IPV6> using ENABLE_IPV6 from appviewx.conf (true or false).
  3. The IU prints the file path and runs a YAML syntax and structure check.
  4. When prompted Open the file in an editor now? (Yes/no):, choose Yes to edit or No to skip. The default is Yes. The editor is $EDITOR, or $VISUAL, or vi. After a successful edit, the IU prints: Edit complete and YAML validated. Use option 5 to apply the configuration.
  5. The editor is $EDITOR, or $VISUAL, or vi.
  6. After the editor exits, YAML is validated again.
  7. Use option 5 to apply. Option 4 does not apply configuration to the cluster.

Required sections

IU validation requires these top-level keys:

node_manager: # host-level checks ; cluster_manager:
# cluster-level checks

Optional/Per-plugin settings

Each plugin under node_manager and cluster_manager can set validate, autofix, command_timeout, and check-specific keys (for example expected_active, service_names, max_usage_percent, namespace). Plugins that are not required in an environment can set validate: false. Do not remove the required top-level sections.
Section Role Notes from current template
node_manager.service Host services (sshd, kubelet, containerd, firewalld, chrony, and others) firewalld expected inactive; several optional services ignore_if_not_found
node_manager.sysctl Kernel parameters including IPv6-related keys ipv6_enabled is set from ENABLE_IPV6 when the file is created from the template
node_manager.dns / ulimit / path / command / packages Resolver, file limits, install paths, identity/SELinux, package presence Keep template structure
node_manager.images / resources / metrics Image presence and host resource thresholds Adjust thresholds to site standards
node_manager.kube_certs kubelet/kubeadm certificate validity autofix is false by default. Enabling it can cause AppViewX downtime.
cluster_manager.node_health Node Ready status label_selector may be empty for all nodes
cluster_manager.pod_health Pod health by namespace (avx, kube-system, istio-system) autofix per namespace as in the template
cluster_manager.basic_checks Redis, Mongo, namespace, and web certificate checks Mongo replica lag keys are present in the template
cluster_manager.metrics_cluster Cluster metrics collection validate-only in the template
YAML validation
  • File must exist, parse as YAML, and be a non-empty mapping.
  • node_manager and cluster_manager must both be present.
  • Invalid YAML is reported immediately. Fix the file and run option 4 again.

Recommended workflow after editing

Prepare / Edit Configuration (option 4), then Validate Configuration (automatic YAML check), then Apply Configuration (option 5), then Verify Configuration (option 3 and kubectl).
Note: Editing YAML does not change the running system until you Apply Configuration. Reload Configuration publishes auto_remediation.conf, not the YAML check file.

Apply Configuration

Purpose

Publish scripts/auto_remediation_config.yaml to Node Manager (host config.yaml and service restart) and/or Cluster Manager (ConfigMap cluster-manager-config and deployment rollout).

Prerequisites
  • Kubernetes API reachable. Namespace auto-remediation exists.
  • auto_remediation_config.yaml exists (created from the template if missing) and passes YAML validation.
  • Deployment auto-remediation-operator exists.
  • For Node Manager: Node Manager installed on target nodes and reachable over SSH.
Configuration validation
  • The IU validates YAML syntax and required top-level sections before prompting for type.
  • The apply script compares desired YAML to the currently applied configuration (trailing whitespace normalized).
Apply procedure
  1. Select 5. Apply Configuration.
  2. Choose the target:
  3. Confirm Apply configuration to <target>? (yes/no):. The default is Yes. <target> is Node Manager, Cluster Manager, or All Components.
  4. The IU runs: auto_remediation.sh --apply-config <node-manager|cluster-manager|all> --yes --non-interactive.
    Note: 1. Node Manager 2. Cluster Manager 3. All Components 4. Back
Kubernetes and host resources involved
Target What is updated
Node Manager Host file /usr/libexec/appviewx/config.yaml (copied via the installer plugin path) and systemctl restart of appviewx-node-manager.service
Cluster Manager ConfigMap cluster-manager-config (config.yaml, MONITORING_INTERVAL, appviewx.conf) in namespace auto-remediation; rollout restart of deployment cluster-manager
All Only the component whose YAML actually changed is redeployed

No-change optimization

If the desired YAML matches what is already applied, redeployment is skipped. The expected message is:
Note:
  • No configuration changes detected. Skipping redeployment.
  • A successful skip is not a failed apply and is not a full apply. Do not restart pods manually after a skip unless you are troubleshooting a mismatch between files on disk and live resources.
  • When type is all, Node Manager and Cluster Manager are evaluated independently. Unchanged sides are skipped with an explicit skip line. Changed sides are applied.

Expected output

Use option 3 to export and review applied YAML.
Situation Expected result
No YAML differences [INFO] No configuration changes detected. Skipping redeployment. Exit code 0.
Changes on selected target Component apply, then [SUCCESS] Configuration applied successfully for type: <type>
User answers no at IU confirm Apply cancelled.
Error handling and troubleshooting
Issue Resolution
Invalid YAML / missing sections Fix auto_remediation_config.yaml (option 4), then retry apply.
Operator not found Install Auto Remediation before apply.
Node Manager not installed or not running on a node Install Node Manager or start appviewx-node-manager.service.
SSH unreachable Fix SSH, PEM, and user as used by appviewx.conf.
ConfigMap apply or rollout timeout Describe and collect logs for cluster-manager, then retry apply after the API is healthy.

Run Alert Cycle

Purpose

Immediately process waiting Auto Remediation alert records (send a digest according to SMTP and alert settings) instead of waiting for ALERT_CRON_SCHEDULE.

What the alert cycle does
  • Checks that the operator is installed and ready.
  • Loads auto_remediation.conf and requires manual trigger to be enabled.
  • Requires complete SMTP (server, from, to, and password or secret) and that alerting is enabled on the operator (CronJob auto-remediation-operator-alert-emailer and smtp-credentials).
  • Triggers processing via a one-off Job from the alert-emailer CronJob (preferred), then in-pod CLI, then API as fallbacks.
  • Does not reinstall components.
Prerequisites
  • Namespace auto-remediation. Operator deployment ready. ConfigMap auto-remediation-operator-config present.
  • SMTP configured and applied (IU option 2). SMTP_SERVER, SMTP_FROM, and SMTP_TO must be set. A password must be available as SMTP_PASSWORD_ENCRYPTED or secret smtp-credentials.
  • ALERT_MANUAL_TRIGGER_ENABLED is true, or unset (the implementation default is true).

ALERT_MANUAL_TRIGGER_ENABLED

This flag exists in auto_remediation.conf-template.template (default true). The Interactive IU and auto_remediation.sh --run-alert-cycle both enforce it. If it is missing from a conf file, the tools treat it as true. If it is set to false, Run Alert Cycle is blocked.
Note: ALERT_MANUAL_TRIGGER_ENABLED=true

If you must enable it after it was set to false: set the value in auto_remediation.conf, then use Reload Configuration (option 1).

Execution procedure
  1. Complete Apply Configuration and SMTP sync first.
  2. Select 6. Run Alert Cycle.
  3. The IU runs auto_remediation.sh --run-alert-cycle.
  4. Review the on-screen summary and logs.

Expected output

On success the IU prints Alert cycle completed. The script logs Manual alert processing completed. If a cycle is already running, the script may report that processing is already in progress (warning; exit 2).

Log locations
Log Location
Installer / IU script log logs/auto_remediation_<timestamp>.log
Operator kubectl logs -n auto-remediation -l app=auto-remediation-operator
Alert emailer Job kubectl get jobs -n auto-remediation ; kubectl logs -n auto-remediation <job-pod>

Validation

Confirm intended recipients received mail only after SMTP is known-good. Do not log passwords while testing SMTP.

kubectl get cronjob auto-remediation-operator-alert-emailer -n auto-remediation ; kubectl get jobs -n auto-remediation
Failure scenarios and troubleshooting
Issue Cause Resolution
SMTP has not been configured Empty SMTP_SERVER / SMTP_FROM / SMTP_TO or missing password Option 2, then retry
ALERT_MANUAL_TRIGGER_ENABLED is not true Flag set to false Set true in auto_remediation.conf, Reload Configuration, retry
Operator not ready Pods crash-looping or missing kubectl get pods -n auto-remediation; repair install
Alert-emailer CronJob missing SMTP/alerting not applied to operator Option 2 (or reload after SMTP is complete)
Already in progress Concurrent cycle Wait; check jobs; retry
No waiting alerts Nothing queued Not a platform failure. Wait for monitoring or reproduce an issue.

Generate Report

Purpose

Generate an Auto Remediation report from the RQLite database (issues and remediation history stored by the platform).

Supported formats

The Interactive IU supports only:
  • HTML
  • CSV
  • Back.
Format Output
html Single file auto_remediation_report_YYYYMMDD_HHMMSS.html
csv Directory auto_remediation_report_YYYYMMDD_HHMMSS/ containing per-table CSV files

PDF, JSON, and other formats are not offered by the current IU and are not supported from this menu.

How to generate a report
  1. Select 7. Generate Report.
  2. Choose 1 for HTML or 2 for CSV.
  3. The IU runs auto_remediation.sh --report --format html or --format csv.

Output directory

Reports are written to REPORT_OUTPUT_DIR from auto_remediation.conf. If that value is empty, the generator uses ./auto_remediation_reports/ relative to the scripts working directory. The shipped template default is /tmp/auto-remediation-reports.

Report naming convention

Prerequisites
  • StatefulSet rqlite or Service rqlite in namespace auto-remediation.
  • Secret rqlite-auth present.
  • Report generator packaged at yaml/appviewx_auto_remediation/rqlite/reports/generate_report.py.
  • Python at <INSTALLATION_PATH>/appviewx_dependencies/appviewx_addons/Python_Linux/bin/python, or python3 on PATH.
  • REPORT_OUTPUT_DIR (or the default directory) is writable.

Command examples (Support)

cd <installer>/appviewx_kubernetes/scripts ; ./auto_remediation.sh --report --format html ; 
./auto_remediation.sh --report --format csv ; ./auto_remediation.sh --report --help

Expected output

The IU prints Report generated successfully. The script prints Report generated successfully. and Report saved to: followed by the absolute output path.

Note: [SUCCESS] Report generated successfully. [SUCCESS] Report saved to: /tmp/auto-remediation-reports/auto_remediation_report_20260824_120000.html

Exit

Option 8 exits the Auto Remediation Configuration submenu and returns to the Interactive IU main menu. It does not uninstall Auto Remediation, does not write configuration, and does not stop Kubernetes workloads. The IU prints Exiting Auto Remediation configuration...

Ctrl+C during a submenu action cancels that operation. Ctrl+C at the submenu prompt also exits the submenu.