Troubleshooting

Primary logs: logs/auto_remediation_*.log and Interactive IU logs. Do not attach secrets.

Configuration validation failure

Item Detail
Symptoms IU or script reports invalid YAML, missing node_manager or cluster_manager, empty required conf keys, or auto_remediation.conf appears incomplete.
Possible cause Edit errors, truncated copy, template not instantiated, or the wrong file was edited.
Resolution Restore structure from auto_remediation_config-template.yaml.template or auto_remediation.conf-template.template. Re-run option 4, then option 5 or 1.
Validation Option 4 reports YAML validated. Reload or apply validation passes.

Apply Configuration failure

Item Detail
Symptoms IU reports Apply Configuration failed with a non-zero exit. Script errors on ConfigMap, SSH, or systemd.
Possible cause Operator missing, Node Manager missing, SSH or sudo failure, or apply error after a real YAML change.
Resolution Fix prerequisites. Check logs/auto_remediation_*.log. Retry option 5. For Node Manager, confirm systemctl status appviewx-node-manager.service.
Validation Skip message if unchanged, or success for type. Pods Ready.

Reload Configuration failure

Item Detail
Symptoms Reload Configuration failed, abort because RQLite nodes changed, or rollback message.
Possible cause Invalid conf, Helm timeout, sudo missing for Node Manager, or RQLite topology edited in conf.
Resolution If RQLite nodes changed, reinstall or redeploy instead of reload. Otherwise fix conf, review backup under logs/auto_remediation_config_backup_*, retry option 1.
Validation [INFO] skip if unchanged, or [SUCCESS] reload. kubectl get pods -n auto-remediation.

SMTP synchronization failure

Item Detail
Symptoms Product SMTP missing or incomplete, sync exit non-zero, or SMTP still incomplete after sync.
Possible cause AppViewX UI SMTP not set, application database unreachable, manual source blocking overwrite, or operator missing.
Resolution Configure Settings -> General -> SMTP. Retry option 2. Do not paste passwords. If manual SMTP is intended, complete it via option 1 conf edit rather than forcing product overwrite.
Validation Masked preview shows server, from, and to. smtp-credentials exists. Option 6 is no longer blocked on SMTP.

Alert Cycle failure

Item Detail
Symptoms Cannot run Alert Cycle; SMTP errors; ALERT_MANUAL_TRIGGER_ENABLED; operator not ready; already in progress.
Possible cause Prerequisites listed in section 5.6 are not met, or a concurrent job is running.
Resolution Complete options 5 and 2. Set ALERT_MANUAL_TRIGGER_ENABLED=true if it is false, then reload. Inspect jobs and logs in namespace auto-remediation.
Validation Alert cycle completed, or a clear no-waiting-alerts or already-in-progress message.

Report generation failure

Item Detail
Symptoms Cannot generate report; generator missing; RQLite not reachable; directory not writable.
Possible cause RQLite not installed, rqlite-auth missing, Python missing, or disk permissions.
Resolution See section 5.7. Repair the RQLite Helm release. Fix REPORT_OUTPUT_DIR permissions.
Validation HTML file or CSV directory created. IU success line.

RQLite connectivity issues

Item Detail
Symptoms Reports fail, operator cannot persist findings, or secret rqlite-auth is missing.
Possible cause StatefulSet not Ready, service rqlite down, auth secret lost, or network policy.
Resolution kubectl get sts,pods,svc -n auto-remediation. Check Helm release rqlite. Restore from install if required.
Validation rqlite pods Ready. Report option 7 succeeds.

Missing Kubernetes resources

Item Detail
Symptoms IU prerequisite errors for namespace, operator, ConfigMaps, secrets, or rqlite.
Possible cause Failed or incomplete AppViewX installation, or a partial uninstall of Auto Remediation.
Resolution Re-run Auto Remediation install from scripts: ./auto_remediation.sh --install --non-interactive. Do not invent resource names.
Validation Resources listed in section 3 are present.

Auto-Fix/Remediation Failures

Item Detail
Symptoms Issues persist in reports or alerts. Pods remain unhealthy. Node Manager checks fail.
Possible cause Plugin autofix is false (including kube_certs by default), check not enabled, insufficient sudo, or unsafe autofix skipped.
Resolution Inspect auto_remediation_config.yaml autofix flags. Do not enable kube_certs autofix unless downtime is accepted. Fix underlying infrastructure, re-apply YAML, and review operator and node-manager logs.
Validation Option 3 export shows intended autofix flags. Subsequent report or alert reflects the new state.