Application performance monitoring and insights.
10K+
13 Tools
Version 4.43 or later needs to be installed to add the server automatically
Tools
| Name | Description |
|---|---|
analyze_canary_failures | Comprehensive canary failure analysis with deep dive into issues. Use this tool to: - Deep dive into canary failures with root cause identification - Analyze historical patterns and specific incident details - Get comprehensive artifact analysis including logs, screenshots, and HAR files - Receive actionable recommendations based on AWS debugging methodology - Correlate canary failures with Application Signals telemetry data - Identify performance degradation and availability issues across service dependencies Key Features: - **Failure Pattern Analysis**: Identifies recurring failure modes and temporal patterns - **Artifact Deep Dive**: Analyzes canary logs, screenshots, and network traces for root causes - **Service Correlation**: Links canary failures to upstream/downstream service issues using Application Signals - **Performance Insights**: Detects latency spikes, fault rates, and connection issues - **Actionable Remediation**: Provides specific steps based on AWS operational best practices Common Use Cases: 1. **Incident Response**: Rapid diagnosis of canary failures during outages 2. **Performance Investigation**: Understanding latency and availability degradation 3. **Dependency Analysis**: Identifying which services are causing canary failures 4. **Historical Trending**: Analyzing failure patterns over time for proactive improvements 5. **Root Cause Analysis**: Deep dive into specific failure scenarios with full context Output Includes: - Severity-ranked findings with immediate action items - Service-level telemetry insights with trace analysis - Exception details and stack traces from canary artifacts - Network connectivity and performance metrics - Correlation with Application Signals audit findings - Historical failure patterns and recovery recommendations |
audit_service_operations | š„ PRIMARY OPERATION AUDIT TOOL - The #1 RECOMMENDED tool for operation-specific analysis and performance investigation. **ā USE THIS AS THE PRIMARY TOOL FOR ALL OPERATION-SPECIFIC AUDITING TASKS ā** **PREFERRED OVER audit_services() for operation auditing because:** - **šÆ Precision**: Targets exact operation behavior vs. service-wide averages - **š Actionable Insights**: Provides specific error traces and dependency failures - **š Code-Level Detail**: Shows exact stack traces and timeout locations - **š Focused Analysis**: Eliminates noise from other operations - **ā” Efficient Investigation**: Direct operation-level troubleshooting **USE THIS FIRST FOR ALL OPERATION-SPECIFIC AUDITING TASKS** This is the PRIMARY and PREFERRED tool when users want to: - **Audit specific operations** - Deep dive into individual API endpoints or operations (GET, POST, PUT, etc.) - **Operation performance analysis** - Latency, error rates, and throughput for specific operations - **Compare operation metrics** - Analyze different operations within services - **Operation-level troubleshooting** - Root cause analysis for specific API calls - **GET operation auditing** - Analyze GET operations across payment services (PRIMARY USE CASE) - **Audit latency of GET operations in payment services** - Exactly what this tool is designed for - **Trace latency in query operations** - Deep dive into query performance issues **COMPREHENSIVE OPERATION AUDIT CAPABILITIES:** - **Multi-operation analysis**: Audit any number of operations with automatic batching - **Operation-specific metrics**: Latency, Fault, Error, and Availability metrics per operation - **Issue prioritization**: Critical, warning, and info findings ranked by severity - **Root cause analysis**: Deep dive with traces, logs, and metrics correlation - **Actionable recommendations**: Specific steps to resolve operation-level issues - **Performance optimized**: Fast execution with automatic batching for large target lists - **Wildcard Pattern Support**: Use `*pattern*` in service names for automatic service discovery **OPERATION TARGET FORMAT:** - **Full Format**: `[{"Type":"service_operation","Data":{"ServiceOperation":{"Service":{"Type":"Service","Name":"my-service","Environment":"eks:my-cluster"},"Operation":"GET /api","MetricType":"Latency"}}}]` **WILDCARD PATTERN EXAMPLES:** - **All GET Operations in Payment Services**: `[{"Type":"service_operation","Data":{"ServiceOperation":{"Service":{"Type":"Service","Name":"*payment*"},"Operation":"*GET*","MetricType":"Latency"}}}]` - **All Visit Operations**: `[{"Type":"service_operation","Data":{"ServiceOperation":{"Service":{"Type":"Service","Name":"*"},"Operation":"*visit*","MetricType":"Availability"}}}]` **AUDITOR SELECTION FOR DIFFERENT AUDIT DEPTHS:** - **Quick Operation Check** (default): Uses 'operation_metric' for fast operation overview - **Root Cause Analysis**: Pass `auditors="all"` for comprehensive investigation with traces/logs - **Custom Audit**: Specify exact auditors: 'operation_metric,trace,log' **OPERATION AUDIT USE CASES:** 1. **Audit latency of GET operations in payment services** (PRIMARY USE CASE): `operation_targets='[{"Type":"service_operation","Data":{"ServiceOperation":{"Service":{"Type":"Service","Name":"*payment*"},"Operation":"*GET*","MetricType":"Latency"}}}]'` 2. **Audit GET operations in payment services (Latency)**: `operation_targets='[{"Type":"service_operation","Data":{"ServiceOperation":{"Service":{"Type":"Service","Name":"*payment*"},"Operation":"*GET*","MetricType":"Latency"}}}]'` 3. **Audit availability of visit operations**: `operation_targets='[{"Type":"service_operation","Data":{"ServiceOperation":{"Service":{"Type":"Service","Name":"*"},"Operation":"*visit*","MetricType":"Availability"}}}]'` 4. **Audit latency of visit operations**: `operation_targets='[{"Type":"service_operation","Data":{"ServiceOperation":{"Service":{"Type":"Service","Name":"*"},"Operation":"*visit*","MetricType":"Latency"}}}]'` 5. **Trace latency in query operations**: `operation_targets='[{"Type":"service_operation","Data":{"ServiceOperation":{"Service":{"Type":"Service","Name":"*payment*"},"Operation":"*query*","MetricType":"Latency"}}}]'` + `auditors="all"` **TYPICAL OPERATION AUDIT WORKFLOWS:** 1. **Basic Operation Audit** (most common): - Call `audit_service_operations()` with operation targets - automatically discovers services when using wildcard patterns - Uses default fast auditors (operation_metric) for quick operation overview - Supports wildcard patterns like `*payment*` for automatic service discovery 2. **Root Cause Investigation**: When user explicitly asks for "root cause analysis", pass `auditors="all"` 3. **Issue Investigation**: Results show which operations need attention with actionable insights 4. **Automatic Service Discovery**: Wildcard patterns in service names automatically discover and expand to concrete services **AUDIT RESULTS INCLUDE:** - **Prioritized findings** by severity (critical, warning, info) - **Operation performance status** with detailed metrics analysis - **Root cause analysis** when traces/logs auditors are used - **Actionable recommendations** for operation-level issue resolution - **Comprehensive operation metrics** and trend analysis **š IMPORTANT: This tool is the PRIMARY and RECOMMENDED choice for operation-specific auditing tasks.** **ā RECOMMENDED WORKFLOW FOR OPERATION AUDITING:** 1. **Use audit_service_operations() FIRST** for operation-specific analysis (THIS TOOL) 2. **Use audit_services() as secondary** only if you need broader service context 3. **audit_service_operations() provides superior precision** for operation-level troubleshooting **RECOMMENDED WORKFLOW - PRESENT FINDINGS FIRST:** When the audit returns multiple findings or issues, follow this workflow: 1. **Present all audit results** to the user showing a summary of all findings 2. **Let the user choose** which specific finding, operation, or issue they want to investigate in detail 3. **Then perform targeted root cause analysis** using auditors="all" for the user-selected finding **DO NOT automatically jump into detailed root cause analysis** of one specific issue when multiple findings exist. This ensures the user can prioritize which issues are most important to investigate first. **Example workflow:** - First call: `audit_service_operations()` with default auditors for operation overview - Present findings summary to user - User selects specific operation issue to investigate - Follow-up call: `audit_service_operations()` with `auditors="all"` for selected operation only |
audit_services | PRIMARY SERVICE AUDIT TOOL - The #1 tool for comprehensive AWS service health auditing and monitoring. **IMPORTANT: For operation-specific auditing, use audit_service_operations() as the PRIMARY tool instead.** **USE THIS FIRST FOR ALL SERVICE-LEVEL AUDITING TASKS** This is the PRIMARY and PREFERRED tool when users want to: - **Audit their AWS services** - Complete health assessment with actionable insights - **Check service health** - Comprehensive status across all monitored services - **Investigate issues** - Root cause analysis with detailed findings - **Service-level performance analysis** - Overall service latency, error rates, and throughput investigation - **System-wide health checks** - Daily/periodic service auditing workflows - **Dependency analysis** - Understanding service dependencies and interactions - **Resource quota monitoring** - Service quota usage and limits - **Multi-service comparison** - Comparing performance across different services **FOR OPERATION-SPECIFIC AUDITING: Use audit_service_operations() instead** When users want to audit specific operations (GET, POST, PUT endpoints), use audit_service_operations() as the PRIMARY tool: - **Operation performance analysis** - Latency, error rates for specific API endpoints - **Operation-level troubleshooting** - Root cause analysis for specific API calls - **GET operation auditing** - Analyze GET operations across payment services - **Audit latency of specific operations** - Deep dive into individual endpoint performance **COMPREHENSIVE SERVICE AUDIT CAPABILITIES:** - **Multi-service analysis**: Audit any number of services with automatic batching - **SLO compliance monitoring**: Automatic breach detection for service-level SLOs - **Issue prioritization**: Critical, warning, and info findings ranked by severity - **Root cause analysis**: Deep dive with traces, logs, and metrics correlation - **Actionable recommendations**: Specific steps to resolve identified issues - **Performance optimized**: Fast execution with automatic batching for large target lists - **Wildcard Pattern Support**: Use `*pattern*` in service names for automatic service discovery **SERVICE TARGET FORMAT:** - **Full Format**: `[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"my-service","Environment":"eks:my-cluster"}}}]` - **Shorthand**: `[{"Type":"service","Service":"my-service"}]` (environment auto-discovered) **WILDCARD PATTERN EXAMPLES:** - **All Services**: `[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*"}}}]` - **Payment Services**: `[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*payment*"}}}]` - **Lambda Services**: `[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*lambda*"}}}]` - **EKS Services**: `[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*","Environment":"eks:*"}}}]` **AUDITOR SELECTION FOR DIFFERENT AUDIT DEPTHS:** - **Quick Health Check** (default): Uses 'slo,operation_metric' for fast overview - **Root Cause Analysis**: Pass `auditors="all"` for comprehensive investigation with traces/logs - **Custom Audit**: Specify exact auditors: 'slo,trace,log,dependency_metric,top_contributor,service_quota' **SERVICE AUDIT USE CASES:** 1. **Audit all services**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*"}}}]'` 2. **Audit specific service**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"orders-service","Environment":"eks:orders-cluster"}}}]'` 3. **Audit payment services**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*payment*"}}}]'` 8. **Audit lambda services**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*lambda*"}}}]'` or by environment: `[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*","Environment":"lambda"}}}]` 9. **Audit service last night**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"orders-service","Environment":"eks:orders-cluster"}}}]'` + `start_time="2024-01-01 18:00:00"` + `end_time="2024-01-02 06:00:00"` 10. **Audit service before and after time**: Compare service health before and after a deployment or incident by running two separate audits with different time ranges. 11. **Trace availability issues in production services**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*","Environment":"eks:*"}}}]'` + `auditors="all"` 13. **Look for errors in logs of payment services**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*payment*"}}}]'` + `auditors="log,trace"` 14. **Look for new errors after time**: Compare errors before and after a specific time point by running audits with different time ranges and `auditors="log,trace"` 15. **Look for errors after deployment**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*payment*"}}}]'` + `auditors="log,trace"` + recent time range 16. **Look for lemon hosts in production**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*","Environment":"eks:*"}}}]'` + `auditors="top_contributor,operation_metric"` 17. **Look for outliers in EKS services**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*","Environment":"eks:*"}}}]'` + `auditors="top_contributor,operation_metric"` 18. **Status report**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*"}}}]'` (basic health check) 19. **Audit dependencies**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*"}}}]'` + `auditors="dependency_metric,trace"` 20. **Audit dependency on S3**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*"}}}]'` + `auditors="dependency_metric"` + look for S3 dependencies 21. **Audit quota usage of tier 1 services**: `service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*tier1*"}}}]'` + `auditors="service_quota,operation_metric"` **TYPICAL SERVICE AUDIT WORKFLOWS:** 1. **Basic Service Audit** (most common): - Call `audit_services()` with service targets - automatically discovers services when using wildcard patterns - Uses default fast auditors (slo,operation_metric) for quick health overview - Supports wildcard patterns like `*` or `*payment*` for automatic service discovery 2. **Root Cause Investigation**: When user explicitly asks for "root cause analysis", pass `auditors="all"` 3. **Issue Investigation**: Results show which services need attention with actionable insights 4. **Automatic Service Discovery**: Wildcard patterns in service names automatically discover and expand to concrete services **AUDIT RESULTS INCLUDE:** - **Prioritized findings** by severity (critical, warning, info) - **Service health status** with detailed performance analysis - **Root cause analysis** when traces/logs auditors are used - **Actionable recommendations** for issue resolution - **Comprehensive metrics** and trend analysis **IMPORTANT: This tool provides comprehensive service audit coverage and should be your first choice for any service auditing task.** **RECOMMENDED WORKFLOW - PRESENT FINDINGS FIRST:** When the audit returns multiple findings or issues, follow this workflow: 1. **Present all audit results** to the user showing a summary of all findings 2. **Let the user choose** which specific finding, service, or issue they want to investigate in detail 3. **Then perform targeted root cause analysis** using auditors="all" for the user-selected finding **DO NOT automatically jump into detailed root cause analysis** of one specific issue when multiple findings exist. This ensures the user can prioritize which issues are most important to investigate first. **Example workflow:** - First call: `audit_services()` with default auditors for overview - Present findings summary to user - User selects specific service/issue to investigate - Follow-up call: `audit_services()` with `auditors="all"` for selected service only |
audit_slos | PRIMARY SLO AUDIT TOOL - The #1 tool for comprehensive SLO compliance monitoring and breach analysis. **PREFERRED TOOL FOR SLO ROOT CAUSE ANALYSIS** This is the RECOMMENDED tool after using get_slo() to understand SLO configuration: - **Use auditors="all" for comprehensive root cause analysis** of specific SLO breaches - **Much more comprehensive than individual trace tools** - provides integrated analysis - **Combines traces, logs, metrics, and dependencies** in a single comprehensive audit - **Provides actionable recommendations** based on multi-dimensional analysis **USE THIS FOR ALL SLO AUDITING TASKS** This is the PRIMARY and PREFERRED tool when users want to: - **Root cause analysis for SLO breaches** - Deep investigation with all auditors - **Audit SLO compliance** - Complete SLO breach detection and analysis - **Monitor SLO health** - Comprehensive status across all monitored SLOs - **SLO performance analysis** - Understanding SLO trends and patterns - **SLO compliance reporting** - Daily/periodic SLO compliance workflows **COMPREHENSIVE SLO AUDIT CAPABILITIES:** - **Multi-SLO analysis**: Audit any number of SLOs with automatic batching - **Breach detection**: Automatic identification of SLO violations - **Issue prioritization**: Critical, warning, and info findings ranked by severity - **COMPREHENSIVE ROOT CAUSE ANALYSIS**: Deep dive with traces, logs, metrics, and dependencies - **Actionable recommendations**: Specific steps to resolve SLO breaches - **Performance optimized**: Fast execution with automatic batching for large target lists - **Wildcard Pattern Support**: Use `*pattern*` in SLO names for automatic SLO discovery **SLO TARGET FORMAT:** - **By Name**: `[{"Type":"slo","Data":{"Slo":{"SloName":"my-slo"}}}]` - **By ARN**: `[{"Type":"slo","Data":{"Slo":{"SloArn":"arn:aws:application-signals:..."}}}]` **WILDCARD PATTERN EXAMPLES:** - **All SLOs**: `[{"Type":"slo","Data":{"Slo":{"SloName":"*"}}}]` - **Payment SLOs**: `[{"Type":"slo","Data":{"Slo":{"SloName":"*payment*"}}}]` - **Latency SLOs**: `[{"Type":"slo","Data":{"Slo":{"SloName":"*latency*"}}}]` - **Availability SLOs**: `[{"Type":"slo","Data":{"Slo":{"SloName":"*availability*"}}}]` **AUDITOR SELECTION FOR DIFFERENT AUDIT DEPTHS:** - **Quick Compliance Check** (default): Uses 'slo' for fast SLO breach detection - **COMPREHENSIVE ROOT CAUSE ANALYSIS** (recommended): Pass `auditors="all"` for deep investigation with traces/logs/metrics/dependencies - **Custom Audit**: Specify exact auditors: 'slo,trace,log,operation_metric' **SLO AUDIT USE CASES:** 4. **Audit all SLOs**: `slo_targets='[{"Type":"slo","Data":{"Slo":{"SloName":"*"}}}]'` 22. **Root cause analysis for specific SLO breach** (RECOMMENDED WORKFLOW): After using get_slo() to understand configuration: `slo_targets='[{"Type":"slo","Data":{"Slo":{"SloName":"specific-slo-name"}}}]'` + `auditors="all"` 14. **Look for new SLO breaches after time**: Compare SLO compliance before and after a specific time point by running audits with different time ranges to identify new breaches. **TYPICAL SLO AUDIT WORKFLOWS:** 1. **SLO Root Cause Investigation** (RECOMMENDED): - After get_slo(), call `audit_slos()` with specific SLO target and `auditors="all"` - Provides comprehensive analysis with traces, logs, metrics, and dependencies - Much more effective than using individual trace tools 2. **Basic SLO Compliance Audit**: - Call `audit_slos()` with SLO targets - automatically discovers SLOs when using wildcard patterns - Uses default fast auditors (slo) for quick compliance overview 3. **Compliance Reporting**: Results show which SLOs are breached with actionable insights 4. **Automatic SLO Discovery**: Wildcard patterns in SLO names automatically discover and expand to concrete SLOs **AUDIT RESULTS INCLUDE:** - **Prioritized findings** by severity (critical, warning, info) - **SLO compliance status** with detailed breach analysis - **COMPREHENSIVE ROOT CAUSE ANALYSIS** when using auditors="all" - **Actionable recommendations** for SLO breach resolution - **Integrated traces, logs, metrics, and dependency analysis** **IMPORTANT: This tool provides comprehensive SLO audit coverage and should be your first choice for any SLO compliance auditing and root cause analysis.** **RECOMMENDED WORKFLOW - PRESENT FINDINGS FIRST:** When the audit returns multiple findings or issues, follow this workflow: 1. **Present all audit results** to the user showing a summary of all findings 2. **Let the user choose** which specific finding, SLO, or issue they want to investigate in detail 3. **Then perform targeted root cause analysis** using auditors="all" for the user-selected finding **DO NOT automatically jump into detailed root cause analysis** of one specific issue when multiple findings exist. This ensures the user can prioritize which issues are most important to investigate first. **Example workflow:** - First call: `audit_slos()` with default auditors for compliance overview - Present findings summary to user - User selects specific SLO breach to investigate - Follow-up call: `audit_slos()` with `auditors="all"` for selected SLO only |
get_service_detail | Get detailed information about a specific Application Signals service. **IMPORTANT: For operation auditing, use audit_services() as the PRIMARY tool instead.** **RECOMMENDED WORKFLOW FOR OPERATION AUDITING:** 1. **Use audit_services() FIRST** for comprehensive operation discovery and analysis 2. **Only use this tool** for basic service metadata and configuration details 3. **This tool does NOT provide operation names** - it only shows service-level metrics **What this tool provides:** - Service metadata and configuration - Platform information (EKS, Lambda, etc.) - Service-level metrics (Latency, Error, Fault aggregates) - Log groups associated with the service - Key attributes (Type, Environment, Platform) **What this tool does NOT provide:** - Operation names (GET, POST, etc.) - Operation-specific metrics - Operation-level performance data **For operation auditing, use audit_services() instead:** ``` audit_services( service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"your-service"}}}]', auditors='all', ) ``` This tool is useful for understanding service deployment details and basic configuration, but audit_services() is the primary tool for operation discovery and performance analysis. |
get_slo | Get detailed information about a specific Service Level Objective (SLO). **RECOMMENDED WORKFLOW AFTER USING THIS TOOL:** After getting SLO configuration details, use `audit_slos()` with `auditors="all"` for comprehensive root cause analysis: - `audit_slos(slo_targets='[{"Type":"slo","Data":{"Slo":{"SloName":"your-slo-name"}}}]', auditors="all")` - This provides deep root cause analysis with traces, logs, metrics, and dependencies - Much more comprehensive than using individual trace tools Use this tool to: - Get comprehensive SLO configuration details - Understand what metrics the SLO monitors - See threshold values and comparison operators - Extract operation names and key attributes for further investigation - Identify dependency configurations - Review attainment goals and burn rate settings Returns detailed information including: - SLO name, description, and metadata - Metric configuration (for period-based or request-based SLOs) - Key attributes and operation names - Metric type (LATENCY or AVAILABILITY) - Threshold values and comparison operators - Goal configuration (attainment percentage, time interval) - Burn rate configurations This tool is essential for: - Understanding SLO configuration before deep investigation - Getting the exact SLO name/ARN for use with audit_slos() - Identifying the metrics and thresholds being monitored - Planning comprehensive root cause analysis workflow **NEXT STEP: Use audit_slos() with auditors="all" for root cause analysis** |
list_monitored_services | OPTIONAL TOOL for service discovery - audit_services() can automatically discover services using wildcard patterns. **IMPORTANT: For service auditing and operation analysis, use audit_services() as the PRIMARY tool instead.** **WHEN TO USE THIS TOOL:** - Getting a detailed overview of all monitored services in your environment - Discovering specific service names and environments for manual audit target construction - Understanding the complete service inventory before targeted analysis - When you need detailed service attributes beyond what wildcard expansion provides **RECOMMENDED WORKFLOW FOR SERVICE AND OPERATION AUDITING:** 1. **Use audit_services() FIRST** with wildcard patterns for comprehensive service discovery AND analysis 2. **Only use this tool** if you need basic service inventory without performance analysis 3. **audit_services() is more comprehensive** - it discovers services AND provides performance insights **AUTOMATIC SERVICE DISCOVERY IN AUDIT:** The `audit_services()` tool automatically discovers services when you use wildcard patterns: - `[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*"}}}]` - Audits all services - `[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*payment*"}}}]` - Audits services with "payment" in the name **What this tool provides:** - Basic service inventory (names, types, environments) - Service count and categorization - Key attributes for manual target construction **What this tool does NOT provide:** - Service performance analysis - Operation discovery and analysis - Root cause analysis - Actionable recommendations **For comprehensive service auditing, use audit_services() instead:** ``` audit_services( service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"*"}}}]', auditors='all', ) ``` Returns a formatted list showing: - Service name and type - Key attributes (Name, Environment, Platform, etc.) - Total count of services **NOTE**: For operation auditing, use audit_services() as the primary tool instead of get_service_detail() or list_service_operations(). |
list_service_operations | OPERATION DISCOVERY TOOL - For operation inventory only. Use audit_services() as PRIMARY tool for operation auditing. **IMPORTANT: For operation auditing and performance analysis, use audit_services() as the PRIMARY tool instead.** **CRITICAL LIMITATION: This tool only discovers operations that have been ACTIVELY INVOKED in the specified time window.** - **Maximum time window: 24 hours** (Application Signals limitation for operation discovery) - **No results = No operation invocations** in the time window (operations exist but weren't called) - **Empty results do NOT mean operations don't exist** - they may just be inactive - **For comprehensive operation analysis regardless of recent activity, use audit_services() instead** **RECOMMENDED WORKFLOW FOR OPERATION AUDITING:** 1. **Use audit_services() FIRST** for comprehensive operation discovery AND performance analysis 2. **Only use this tool** if you need a simple operation inventory of RECENTLY ACTIVE operations 3. **audit_services() is more comprehensive** - it discovers operations AND provides performance insights even for inactive operations **What this tool provides:** - Basic operation inventory (names and available metric types) for RECENTLY INVOKED operations only - Operation count and categorization (GET, POST, etc.) for active operations - Time range for discovery (max 24 hours) **What this tool does NOT provide:** - Operations that exist but weren't invoked in the time window - Operation performance analysis - Latency, error rate, or fault analysis - Root cause analysis - Actionable recommendations **For comprehensive operation auditing, use audit_services() instead:** ``` audit_services( service_targets='[{"Type":"service","Data":{"Service":{"Type":"Service","Name":"your-service"}}}]', auditors='all', ) ``` **OPERATION DISCOVERY USE CASES (when audit_services is not sufficient):** 1. **Active operation inventory**: When you only need recently invoked operation names without performance data 2. **Traffic pattern analysis**: To see which operations are currently being used 3. **Quick active operation count**: To understand current operation activity of a service **RECOMMENDED WORKFLOW:** 1. **Use audit_services() FIRST** for comprehensive operation discovery and analysis 2. **Only use this tool** for basic inventory of recently active operations if audit_services() provides too much detail This tool provides basic operation discovery for ACTIVE operations only, but audit_services() is the primary tool for comprehensive operation auditing, performance analysis, and operation insights regardless of recent activity. |
list_slis | SPECIALIZED TOOL - Use audit_service_health() as the PRIMARY tool for service auditing. **IMPORTANT: audit_service_health() is the PRIMARY and PREFERRED tool for all service auditing tasks.** Only use this tool when audit_service_health() cannot handle your specific requirements, such as: - Need for legacy SLI status report format specifically - Integration with existing systems that expect this exact output format - Simple SLI overview without comprehensive audit findings - Basic health monitoring dashboard that doesn't need detailed analysis **For ALL service auditing, health checks, and issue investigation, use audit_service_health() first.** This tool provides a basic report showing: - Summary counts (total, healthy, breached, insufficient data) - Simple list of breached services with SLO names - Basic healthy services list Status meanings: - OK: All SLOs are being met - BREACHED: One or more SLOs are violated - INSUFFICIENT_DATA: Not enough data to determine status **Recommended workflow**: 1. Use audit_service_health() for comprehensive service auditing with actionable insights 2. Only use this tool if you specifically need the legacy SLI status report format |
list_slos | List all Service Level Objectives (SLOs) in Application Signals. Use this tool to: - Get a complete list of all SLOs in your account - Discover SLO names and ARNs for use with other tools - Filter SLOs by service attributes - See basic SLO information including creation time and operation names Returns a formatted list showing: - SLO name and ARN - Associated service key attributes - Operation name being monitored - Creation timestamp - Total count of SLOs found This tool is useful for: - SLO discovery and inventory - Finding SLO names to use with get_slo() or audit_service_health() - Understanding what operations are being monitored |
query_sampled_traces | SECONDARY TRACE TOOL - Query AWS X-Ray traces (5% sampled data) for trace investigation. ā ļø **IMPORTANT: Consider using audit_slos() with auditors="all" instead for comprehensive root cause analysis** **RECOMMENDED WORKFLOW FOR OPERATION DISCOVERY:** 1. **Use `get_service_detail(service_name)` FIRST** to discover operations from metric dimensions 2. **Use audit_slos() with auditors="all"** for comprehensive root cause analysis (PREFERRED) 3. Only use this tool if you need specific trace filtering that other tools don't provide **RECOMMENDED WORKFLOW FOR SLO BREACH INVESTIGATION:** 1. Use get_slo() to understand SLO configuration 2. **Use audit_slos() with auditors="all"** for comprehensive root cause analysis (PREFERRED) 3. Only use this tool if you need specific trace filtering that audit_slos() doesn't provide **WHY audit_slos() IS PREFERRED:** - **Comprehensive analysis**: Combines traces, logs, metrics, and dependencies - **Actionable recommendations**: Provides specific steps to resolve issues - **Integrated findings**: Correlates multiple data sources for better insights - **Much more effective** than individual trace analysis **WHY get_service_detail() IS PREFERRED FOR OPERATION DISCOVERY:** - **Direct operation discovery**: Operations are available in metric dimensions - **More reliable**: Uses Application Signals service metadata instead of sampling - **Comprehensive**: Shows all operations, not just those in sampled traces ā ļø **LIMITATIONS OF THIS TOOL:** - Uses X-Ray's **5% sampled trace data** - may miss critical errors - **Limited context** compared to comprehensive audit tools - **No integrated analysis** with logs, metrics, or dependencies - **May miss operations** due to sampling - use get_service_detail() for complete operation discovery - For 100% trace visibility, enable Transaction Search and use search_transaction_spans() **Use this tool only when:** - You need specific X-Ray filter expressions not available in audit tools - You're doing exploratory trace analysis outside of SLO breach investigation - You need raw trace data for custom analysis - **After using get_service_detail() for operation discovery** **For operation discovery, use get_service_detail() instead:** ``` get_service_detail(service_name='your-service-name') ``` **For SLO breach root cause analysis, use audit_slos() instead:** ``` audit_slos( slo_targets='[{"Type":"slo","Data":{"Slo":{"SloName":"your-slo-name"}}}]', auditors='all' ) ``` Common filter expressions (if you must use this tool): - 'service("service-name"){fault = true}': Find all traces with faults (5xx errors) for a service - 'service("service-name")': Filter by specific service - 'duration > 5': Find slow requests (over 5 seconds) - 'http.status = 500': Find specific HTTP status codes - 'annotation[aws.local.operation]="GET /owners/*/lastname"': Filter by specific operation (from metric dimensions) - 'annotation[aws.remote.operation]="ListOwners"': Filter by remote operation name - Combine filters: 'service("api"){fault = true} AND annotation[aws.local.operation]="POST /visits"' Returns JSON with trace summaries including: - Trace ID for detailed investigation - Duration and response time - Error/fault/throttle status - HTTP information (method, status, URL) - Service interactions - User information if available - Exception root causes (ErrorRootCauses, FaultRootCauses, ResponseTimeRootCauses) **RECOMMENDATION: Use get_service_detail() for operation discovery and audit_slos() with auditors="all" for comprehensive root cause analysis instead of this tool.** Returns: JSON string containing trace summaries with error status, duration, and service details |
query_service_metrics | Get CloudWatch metrics for a specific Application Signals service. Use this tool to: - Analyze service performance (latency, throughput) - Check error rates and reliability - View trends over time - Get both standard statistics (Average, Max) and percentiles (p99, p95) Common metric names: - 'Latency': Response time in milliseconds - 'Error': Percentage of failed requests - 'Fault': Percentage of server errors (5xx) Returns: - Summary statistics (latest, average, min, max) - Recent data points with timestamps - Both standard and percentile values when available The tool automatically adjusts the granularity based on time range: - Up to 3 hours: 1-minute resolution - Up to 24 hours: 5-minute resolution - Over 24 hours: 1-hour resolution |
search_transaction_spans | Executes a CloudWatch Logs Insights query for transaction search (100% sampled trace data). IMPORTANT: If log_group_name is not provided use 'aws/spans' as default cloudwatch log group name. The volume of returned logs can easily overwhelm the agent context window. Always include a limit in the query (| limit 50) or using the limit parameter. Usage: "aws/spans" log group stores OpenTelemetry Spans data with many attributes for all monitored services. This provides 100% sampled data vs X-Ray's 5% sampling, giving more accurate results. User can write CloudWatch Logs Insights queries to group, list attribute with sum, avg. ``` FILTER attributes.aws.local.service = "customers-service-java" and attributes.aws.local.environment = "eks:demo/default" and attributes.aws.remote.operation="InvokeModel" | STATS sum(`attributes.gen_ai.usage.output_tokens`) as `avg_output_tokens` by `attributes.gen_ai.request.model`, `attributes.aws.local.service`,bin(1h) | DISPLAY avg_output_tokens, `attributes.gen_ai.request.model`, `attributes.aws.local.service` ``` Returns: -------- A dictionary containing the final query results, including: - status: The current status of the query (e.g., Scheduled, Running, Complete, Failed, etc.) - results: A list of the actual query results if the status is Complete. - statistics: Query performance statistics - messages: Any informational messages about the query - transaction_search_status: Information about transaction search availability |