grafana

mirror of https://github.com/grafana/grafana.git synced 2025-02-11 16:15:42 -06:00

Author	SHA1	Message	Date
George Robinson	05d858635c	Alerting: Add metric for inhibition rules (#81119 ) This commit adds a metric for the number of inhibition rules. It matches the metric added upstream in #3681.	2024-01-23 19:43:17 +00:00
Jean-Philippe Quéméner	aa25776f81	Alerting: Add a feature flag to periodically save states (#80987 )	2024-01-23 17:03:30 +01:00
George Robinson	85b9edcd28	Alerting: Fix incorrect initialization of logger (#81099 )	2024-01-23 17:29:38 +02:00
Marcus Efraimsson	6768c6c059	Chore: Remove public vars in setting package (#81018 ) Removes the public variable setting.SecretKey plus some other ones. Introduces some new functions for creating setting.Cfg.	2024-01-23 12:36:22 +01:00
Jean-Philippe Quéméner	eb7e1216a1	feat(alerting): add async state persister (#80763 )	2024-01-22 13:07:11 +01:00
Julien Duchesne	40312c527b	ngalert openapi: Fix ObjectMatchers definition (#79477 ) These don't get marshalled and unmarshalled in the same way as they are represented in Go This PR changes the OpenAPI spec to reflect what the API accepts and sends back	2024-01-19 14:37:11 -05:00
Alexander Weaver	18b9c8fd5f	Alerting: Nilcheck JitterStrategyFrom so it can be used in contexts without feature toggles (#80841 ) Nilcheck so tests can have a nil feature toggles	2024-01-18 15:43:41 -06:00
Alexander Weaver	00a260effa	Alerting: Add setting to distribute rule group evaluations over time (#80766 ) * Simple, per-base-interval jitter * Add log just for test purposes * Add strategy approach, allow choosing between group or rule * Add flag to jitter rules * Add second toggle for jittering within a group * Wire up toggles to strategy * Slightly improve comment ordering * Add tests for offset generation * Rename JitterStrategyFrom * Improve debug log message * Use grafana SDK labels rather than prometheus labels	2024-01-18 12:48:11 -06:00
Julien Duchesne	c9211fbd69	ngalert openapi: Use same `basePath` as rest of Grafana (#79025 ) * ngalert openapi: Use same `basePath` as rest of Grafana Currently, there are two issues that prevent easily merging `ngalert` and grafana openapi specs: - The basePath is different. `grafana` has `/api` and `ngalert` has `/api/v1`. I changed `ngalert` to use `/api` - The `ngalert` endpoints have their basePath in the each operation path. The basePath should actually be omitted --------- Co-authored-by: Yuriy Tseretyan <yuriy.tseretyan@grafana.com>	2024-01-17 11:53:16 -05:00
Jean-Philippe Quéméner	82638d059f	feat(alerting): add state persister interface (#80384 )	2024-01-17 13:33:13 +01:00
Santiago	3217a0dc05	Alerting: Fix state sync errors counter increment (#80702 )	2024-01-17 11:04:27 +01:00
Sofia Papagiannaki	d1dab5828d	Alerting: Update rule API to address folders by UID (#74600 ) * Change ruler API to expect the folder UID as namespace * Update example requests * Fix tests * Update swagger * Modify FIle field in /api/prometheus/grafana/api/v1/rules * Fix ruler export * Modify folder in responses to be formatted as <parent UID>/<title> * Add alerting test with nested folders * Apply suggestion from code review * Alerting: use folder UID instead of title in rule API (#77166) Co-authored-by: Sonia Aguilar <soniaaguilarpeiron@gmail.com> * Drop a few more latent uses of namespace_id * move getNamespaceKey to models package * switch GetAlertRulesForScheduling to use folder table * update GetAlertRulesForScheduling to return folder titles in format `parent_uid/title`. * fi tests * add tests for GetAlertRulesForScheduling when parent uid * fix integration tests after merge * fix test after merge * change format of the namespace to JSON array this is needed for forward compatibility, when we migrate to full paths * update EF code to decode nested folder --------- Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com> Co-authored-by: Virginia Cepeda <virginia.cepeda@grafana.com> Co-authored-by: Sonia Aguilar <soniaaguilarpeiron@gmail.com> Co-authored-by: Alex Weaver <weaver.alex.d@gmail.com> Co-authored-by: Gilles De Mey <gilles.de.mey@gmail.com>	2024-01-17 11:07:39 +02:00
Alexander Weaver	3c796ecc8f	Alerting: Add metric counting rule groups per org (#80669 ) * Refactor, fix bad map hint * Count groups per org	2024-01-16 16:35:56 -06:00
Santiago	3afd94185c	Alerting: Add metric to check for default AM configurations (#80225 ) * Alerting: Add metric to check for default AM configurations * Use a gauge for the config hash * don't go out of bounds when converting uint64 to float64 * expose metric for config hash * update metrics after applying config	2024-01-16 17:12:24 +01:00
Yuri Tseretyan	4b071f5452	Alerting: Fix MuteTiming Get API to return provenance status (#80494 )	2024-01-13 00:16:54 +02:00
Julien Duchesne	2fb03dfa56	fix(swagger): Mute Timing PUT OK status is 202 (#80459 )	2024-01-12 16:58:20 -05:00
Yuri Tseretyan	4479e7218d	Alerting: MuteTiming service return errutil + GetTiming by name (#79772 ) * add get mute timing by name to MuteTimingService * update get mute timing request handler to use the service method * replace validation, uniqueness and used errors with errutils * update mute timing methods return errutil responses * use the term "time interval" in errors bevause mute timings are deprecated in Alertmanager and will be replaced by time intervals in the future. * update create and update methods to return struct instead of pointer	2024-01-12 21:23:44 +02:00
idafurjes	cb419e799b	Remove folderid service test (#80433 ) * Remove FolderID from service tests * Add models * Add folderID pack to publicdashboard tests * Remove folderID from dashboard tests * Remove folderID from folders * Remove folderID from ngalert tests * Remove nolint comment * Add back some tests after rebase	2024-01-12 16:43:39 +01:00
Yuri Tseretyan	77db6a9ca4	Alerting: Fix GetAlertRulesForScheduling to use folder table and join by org_id (#80330 )	2024-01-11 09:21:03 -05:00
Santiago	6c87d9a1e7	Alerting: Stop retries on 4xx status code responses (remote Alertmanager readiness check) (#80350 )	2024-01-11 12:12:35 +01:00
William Wernert	48b5ac779b	Alerting/Annotations: Add annotation backend for Loki alert state history (#78156 ) * Move scope type vars to testutil package * Expose parts of state historian for use in annotation backend * Implement Loki ASH Annotation store This store will only implement the `Get` method of a RepositoryImpl since alert state history writes to Loki elsewhere. * Use interface for Loki HTTP Client * Add tests for Loki ASH Annotation store * Add missing test * Fix lint * Organize tests * Add filter tests * Improve tests * Move filter logic into outer function * Fix lint * Add comment * Fix tests * Fix lint * Rename historian store + refactor * Cleanup historian store * Fix tests * Minor cleanup * Use new `ShouldRecordAnnotation` filter * Fix logic and add tests for this check * Fix typos, remove unused variables, `< 1` -> `== 0` * More closely mimic RBAC filter from xorm to ensure correct logic * Move off weaveworks client * Address PR comments	2024-01-10 18:42:35 -05:00
Matthew Jacobson	afa33f12b2	Alerting: Create alertingQueryOptimization feature flag for alert query optimization (#78932 ) * Alerting: Create feature flag for alert query optimization Adds a feature flag alertingQueryOptimization for an already existing functionality: alert query optimization. This feature flag will now be disabled by default.	2024-01-10 15:52:58 -05:00
Matthew Jacobson	f365d35cf8	Alerting: Show warning when query optimized (#78751 ) * Alerting: Show warning when query optimized * Use frame.AppendNotices * Improve warning to include why and a prompt for action	2024-01-10 14:40:00 -05:00
Santiago	9e78faa7ba	Alerting: Add metrics to the remote Alertmanager struct (#79835 ) * Alerting: Add metrics to the remote Alertmanager struct * rephrase http_requests_failed description * make linter happy * remove unnecessary metrics * extract timed client to separate package * use histogram collector from dskit * remove weaveworks dependency * capture metrics for all requests to the remote Alertmanager (both clients) * use the timed client in the MimirAuthRoundTripper * HTTPRequestsDuration -> HTTPRequestDuration, clean up mimir client factory function * refactor * less git diff * gauge for last readiness check in seconds * initialize LastReadinesCheck to 0, tweak metric names and descriptions * add counters for sync attempts/errors * last config sync and last state sync timestamps (gauges) * change latency metric name * metric for remote Alertmanager mode * code review comments * move label constants to metrics package	2024-01-10 11:18:24 +01:00
Matthew Jacobson	1d4419fbe4	Alerting: Fix NoData & Error alerts not resolving when rule is reset (#80184 ) * Alerting: Fix NoData & Error alerts not resolving when rule is reset On rule reset, when creating the PostableAlerts StateToPostableAlert did not attach the correct NoData/Error alertname and rulename labels to expire/resolve the active alerts when the previous cached state was NoData/Error.	2024-01-09 14:47:19 -05:00
Alexander Weaver	542741f748	Alerting: Log scheduler maxAttempts, guard against invalid retry counts, log retry errors (#80234 ) * Log maxAttempts, add guard, log retry errors * fix whitespace * Initialize evaluator in TestProcessTicks	2024-01-09 13:19:37 -06:00
Matthew Jacobson	aa03b8f8a7	Alerting: Guided legacy alerting upgrade dry-run (#80071 ) This PR has two steps that together create a functional dry-run capability for the migration. By enabling the feature flag alertingPreviewUpgrade when on legacy alerting it will: a. Allow all Grafana Alerting background services except for the scheduler to start (multiorg alertmanager, state manager, routes, …). b. Allow the UI to show Grafana Alerting pages alongside legacy ones (with appropriate in-app warnings that UA is not actually running). c. Show a new “Alerting Upgrade” page and register associated /api/v1/upgrade endpoints that will allow the user to upgrade their organization live without restart and present a summary of the upgrade in a table.	2024-01-05 18:19:12 -05:00
Yuri Tseretyan	72182e02a4	Alerting: Mute timing service tests (#79817 ) split tests for mute timing service to functions for each method this makes it clear the scope of tests	2024-01-06 00:26:15 +02:00
Yuri Tseretyan	494f36e0bd	Alerting: Update provisioning services that handle Alertmanager configuraiton to access config via storage (#79814 ) * extract get and save operations to a alertmanagerConfigStore. this removes duplicated code in service (currently only mute timings) and improves testing * replace generic errors with errutils one with better messages. * update provisioning services to use new store --------- Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com>	2024-01-05 16:15:18 -05:00
Alexander Weaver	a8fb01a502	Swap weaveworks/common utilities for equivalents in grafana/dskit (#80051 ) * Replace histogram collector and grpc injectors * Extract request timing utility * Also vendor test file * Suppress erroneous linter warn	2024-01-05 10:08:38 -06:00
Matthew Jacobson	3537c5440f	Alerting: Refactor migration to return pairs of legacy and upgraded structs (#79719 ) Some refactoring that will simplify next changes for dry-run PRs. This should be no-op as far as the created ngalert resources and database state, though it does change some logs. The key change here is to modify migrateOrg to return pairs of legacy struct + ngalert struct instead of actually persisting the alerts and alertmanager config. This will allow us to capture error information during dry-run migration. It also moves most persistence-related operations such as title deduplication and folder creation to the right before we persist. This will simplify eventual partial migrations (individual alerts, dashboards, channels, ...). Additionally it changes channel code to deal with PostableGrafanaReceiver instead of PostableApiReceiver (integration instead of contact point).	2024-01-05 05:37:13 -05:00
Santiago	1f6575e65e	Alerting: Test MOA in remote secondary mode (#79828 )	2024-01-05 11:05:27 +01:00
Alexander Weaver	90d4704cd7	Alerting: Fix URL timestamp conversion in historian API in annotation mode (#80026 ) Fix timestamp conversion when calling annotation store	2024-01-04 12:40:21 -06:00
Yuri Tseretyan	f6a46744a6	Alerting: Support hysteresis command expression (#75189 ) Backend: * Update the Grafana Alerting engine to provide feedback to HysteresisCommand. The feedback information is stored in state.Manager as a fingerprint of each state. The fingerprint is persisted to the database. Only fingerprints that belong to Pending and Alerting states are considered as "loaded" and provided back to the command. - add ResultFingerprint to state.State. It's different from other fingerprints we store in the state because it is calculated from the result labels. - add rule_fingerprint column to alert_instance - update alerting evaluator to accept AlertingResultsReader via context, and update scheduler to provide it. - add AlertingResultsFromRuleState that implements the new interface in eval package - update getExprRequest to patch the hysteresis command. * Only one "Recovery Threshold" query is allowed to be used in the alert rule and it must be the Condition. Frontend: * Add hysteresis option to Threshold in UI. It's called "Recovery Threshold" * Add test for getUnloadEvaluatorTypeFromCondition * Hide hysteresis in panel expressions * Refactor isInvalid and add test for it * Remove unnecesary React.memo * Add tests for updateEvaluatorConditions --------- Co-authored-by: Sonia Aguilar <soniaaguilarpeiron@gmail.com>	2024-01-04 11:47:13 -05:00
Santiago	a77ba40ed4	Alerting: Use the forked Alertmanager for remote secondary mode (#79646 ) * (WIP) Alerting: Use the forked Alertmanager for remote secondary mode * fall back to using internal AM in case of error * remove TODOs, clean up .ini file, add orgId as part of remote AM config struct * log warnings and errors, fall back to remoteSecondary, fall back to internal AM only * extract logic to decide remote Alertmanager mode to a separate function, switch on mode * tests * make linter happy * remove func to decide remote Alertmanager mode * refactor factory function and options * add default case to switch statement * remove ineffectual assignment	2023-12-21 15:26:31 +01:00
Santiago	c46da8ea9b	Alerting: Update alerting package and imports from cluster and clusterpb (#79786 ) * Alerting: Update alerting package * update to latest commit * alias for imports	2023-12-21 12:34:48 +01:00
Matthew Jacobson	0424d44b39	Alerting: In migration, create one label per channel (#76527 ) * In migration, create one label per channel This PR changes how routing is done by the legacy alerting migration. Previously, we created a single label on each alert rule that contained an array of contact point names. Ex: __contact__="slack legacy testing","slack legacy testing2" This label was then routed against a series of regex-matching policies with continue=true. Ex: __contacts__ =~ ."slack legacy testing". In the case of many contact points, this array could quickly become difficult to manage and difficult to grok at-a-glance. This PR replaces the single __contact__ label with multiple __legacy_c_{contactname}__ labels and simple equality-matching policies. These channel-specific policies are nested in a single route under the top-level route which matches against __legacy_use_channels__ = true for ease of organization. This should improve the experience for users wanting to keep the default migrated routing strategy but who also want to modify which contact points an alert sends to.	2023-12-19 13:25:13 -05:00
Santiago	9945514baa	Alerting: Validate configuration for the remote Alertmanager struct (#79691 ) * Alerting: Validate configuration for the remote Alertmanager struct * add TenantID to test * add OrgID to config struct in tests	2023-12-19 18:41:48 +01:00
Alexander Weaver	65ecde6eed	Alerting: Don't record annotations for mapped NoData transitions, when NoData is mapped to OK (#77164 ) * Exclude mapped nodata transitions when nodata mapped to OK * Fix processEvalResults test * Don't check NoDataState when filtering transition * Add comment to explain purpose of separate function --------- Co-authored-by: William Wernert <william.wernert@grafana.com>	2023-12-18 16:59:32 -05:00
Santiago	f7248efff5	Alerting: Fix panic when creating a new Alertmanager returns an error (#79641 ) Alerting: Fix panic after error creating new Alertmanager	2023-12-18 15:33:07 +01:00
Alexander Weaver	cf8e8852c3	Alerting: Drop NamespaceID from responses on unstable ngalert API endpoints in favor of NamespaceUID (#79359 ) * Drop from API response * Drop from swagger docs * Drop from integration tests * regenerate public swagger docs * Drop from frontend * Drop asserts for namespaceID field	2023-12-15 11:06:53 -06:00
William Wernert	9171bf92bb	Alerting: Add rule ID and title to alert state history Loki entry (#79481 ) * Add rule ID and title to Loki entry * Combine related tests	2023-12-14 13:06:23 -05:00
Santiago	23b4568597	Alerting: Send configuration and state to the remote Alertmanager on shutdown (#78682 ) * Alerting: Send configuration and state to the remote Alertmanager on shutdown * Alerting: Add a sync interval for ApplyConfig in remote secondary mode * add routine to sync states and configs * pass a cancellable context to syncRoutine(), remove tests for ApplyConfig, cache last config in memory * extract logic to update config and state in the remote Alertmanager * get latest config from the database * avoid using separate goroutine for updating state and config * clean up PR * refactor, comments, tests * update tests * remove canceled context from calls to StopAndWait() * create context with timeout and send config and state to remote Alertmanager * update tests * address code review comments	2023-12-13 22:53:09 +01:00
Julien Duchesne	884e0427e6	ngalert openapi: Add `X-Disable-Provenance` to missing operations (#79278 ) Swagger(ngalert): Add `X-Disable-Provenance` to missing operations I added all functions that call the `determineProvenance` function Schema changes are from: `make` in `pkg/services/ngalert/api/tooling` `make swagger-clean && make openapi3-gen` in root	2023-12-13 10:55:59 -05:00
Santiago	91836e7832	Alerting: Add time-based convergence in remote secondary mode (#78809 ) * Alerting: Add a sync interval for ApplyConfig in remote secondary mode * add routine to sync states and configs * pass a cancellable context to syncRoutine(), remove tests for ApplyConfig, cache last config in memory * extract logic to update config and state in the remote Alertmanager * get latest config from the database * avoid using separate goroutine for updating state and config * clean up PR * refactor, comments, tests * update tests * add config struct for remote secondary forked Alertmanager * use errgroups for sync operations * use waitgroup instead of errgroup * remove helper method to sync AMs * check for errors instead of bool syncErr	2023-12-13 13:36:17 +01:00
William Wernert	62bdbe5b44	Annotations/Alerting: Add Loki historian store stub (#78363 ) * Add Loki historian store stub * Add composite store * Use composite store if Loki historian enabled * Split store interface into read/write * Make composite + historian stores read only * Use variadic constructor for composite * Modify Loki store enable logic * Use dskit.concurrency.ForEachJob for parallelism	2023-12-12 17:43:09 -05:00
Alexander Weaver	aa63e91a43	Alerting: Use mux router to match hooks, add support for path variables and methods (#79345 ) * Use a router inside hooks rather than plain string matching * Add test for mismatched method	2023-12-12 14:43:11 -06:00
Julien Duchesne	f977e3faf5	ngalert swagger: Fix status code (#79415 ) This endpoint returns a 202, not a 204 Let me know if we should instead change the response of the API	2023-12-12 13:40:36 -05:00
Santiago	1a5c2cb55b	Alerting: Check whether the internal Alertmanager is ready in remote secondary mode (#79406 ) Alerting: Check whether the internal Alertmanager is ready in remote secondary	2023-12-12 18:33:11 +01:00
gotjosh	cc3c0a2cc2	Alerting: Refactor readiness check (#78799 ) * Alerting: Refactor readiness check Moves the readiness check to the mimir client and removes the need to assert that we have senders - it already has a queue and can hold notifications until we're ready to send them. --------- Signed-off-by: gotjosh <josue.abreu@gmail.com>	2023-12-12 15:34:54 +00:00

1 2 3 4 5 ...

1267 Commits