grafana

mirror of https://github.com/grafana/grafana.git synced 2025-02-25 18:55:37 -06:00

Author	SHA1	Message	Date
Julien Duchesne	2fb03dfa56	fix(swagger): Mute Timing PUT OK status is 202 (#80459 )	2024-01-12 16:58:20 -05:00
Yuri Tseretyan	4479e7218d	Alerting: MuteTiming service return errutil + GetTiming by name (#79772 ) * add get mute timing by name to MuteTimingService * update get mute timing request handler to use the service method * replace validation, uniqueness and used errors with errutils * update mute timing methods return errutil responses * use the term "time interval" in errors bevause mute timings are deprecated in Alertmanager and will be replaced by time intervals in the future. * update create and update methods to return struct instead of pointer	2024-01-12 21:23:44 +02:00
idafurjes	cb419e799b	Remove folderid service test (#80433 ) * Remove FolderID from service tests * Add models * Add folderID pack to publicdashboard tests * Remove folderID from dashboard tests * Remove folderID from folders * Remove folderID from ngalert tests * Remove nolint comment * Add back some tests after rebase	2024-01-12 16:43:39 +01:00
Yuri Tseretyan	77db6a9ca4	Alerting: Fix GetAlertRulesForScheduling to use folder table and join by org_id (#80330 )	2024-01-11 09:21:03 -05:00
Santiago	6c87d9a1e7	Alerting: Stop retries on 4xx status code responses (remote Alertmanager readiness check) (#80350 )	2024-01-11 12:12:35 +01:00
William Wernert	48b5ac779b	Alerting/Annotations: Add annotation backend for Loki alert state history (#78156 ) * Move scope type vars to testutil package * Expose parts of state historian for use in annotation backend * Implement Loki ASH Annotation store This store will only implement the `Get` method of a RepositoryImpl since alert state history writes to Loki elsewhere. * Use interface for Loki HTTP Client * Add tests for Loki ASH Annotation store * Add missing test * Fix lint * Organize tests * Add filter tests * Improve tests * Move filter logic into outer function * Fix lint * Add comment * Fix tests * Fix lint * Rename historian store + refactor * Cleanup historian store * Fix tests * Minor cleanup * Use new `ShouldRecordAnnotation` filter * Fix logic and add tests for this check * Fix typos, remove unused variables, `< 1` -> `== 0` * More closely mimic RBAC filter from xorm to ensure correct logic * Move off weaveworks client * Address PR comments	2024-01-10 18:42:35 -05:00
Matthew Jacobson	afa33f12b2	Alerting: Create alertingQueryOptimization feature flag for alert query optimization (#78932 ) * Alerting: Create feature flag for alert query optimization Adds a feature flag alertingQueryOptimization for an already existing functionality: alert query optimization. This feature flag will now be disabled by default.	2024-01-10 15:52:58 -05:00
Matthew Jacobson	f365d35cf8	Alerting: Show warning when query optimized (#78751 ) * Alerting: Show warning when query optimized * Use frame.AppendNotices * Improve warning to include why and a prompt for action	2024-01-10 14:40:00 -05:00
Santiago	9e78faa7ba	Alerting: Add metrics to the remote Alertmanager struct (#79835 ) * Alerting: Add metrics to the remote Alertmanager struct * rephrase http_requests_failed description * make linter happy * remove unnecessary metrics * extract timed client to separate package * use histogram collector from dskit * remove weaveworks dependency * capture metrics for all requests to the remote Alertmanager (both clients) * use the timed client in the MimirAuthRoundTripper * HTTPRequestsDuration -> HTTPRequestDuration, clean up mimir client factory function * refactor * less git diff * gauge for last readiness check in seconds * initialize LastReadinesCheck to 0, tweak metric names and descriptions * add counters for sync attempts/errors * last config sync and last state sync timestamps (gauges) * change latency metric name * metric for remote Alertmanager mode * code review comments * move label constants to metrics package	2024-01-10 11:18:24 +01:00
Matthew Jacobson	1d4419fbe4	Alerting: Fix NoData & Error alerts not resolving when rule is reset (#80184 ) * Alerting: Fix NoData & Error alerts not resolving when rule is reset On rule reset, when creating the PostableAlerts StateToPostableAlert did not attach the correct NoData/Error alertname and rulename labels to expire/resolve the active alerts when the previous cached state was NoData/Error.	2024-01-09 14:47:19 -05:00
Alexander Weaver	542741f748	Alerting: Log scheduler maxAttempts, guard against invalid retry counts, log retry errors (#80234 ) * Log maxAttempts, add guard, log retry errors * fix whitespace * Initialize evaluator in TestProcessTicks	2024-01-09 13:19:37 -06:00
Matthew Jacobson	aa03b8f8a7	Alerting: Guided legacy alerting upgrade dry-run (#80071 ) This PR has two steps that together create a functional dry-run capability for the migration. By enabling the feature flag alertingPreviewUpgrade when on legacy alerting it will: a. Allow all Grafana Alerting background services except for the scheduler to start (multiorg alertmanager, state manager, routes, …). b. Allow the UI to show Grafana Alerting pages alongside legacy ones (with appropriate in-app warnings that UA is not actually running). c. Show a new “Alerting Upgrade” page and register associated /api/v1/upgrade endpoints that will allow the user to upgrade their organization live without restart and present a summary of the upgrade in a table.	2024-01-05 18:19:12 -05:00
Yuri Tseretyan	72182e02a4	Alerting: Mute timing service tests (#79817 ) split tests for mute timing service to functions for each method this makes it clear the scope of tests	2024-01-06 00:26:15 +02:00
Yuri Tseretyan	494f36e0bd	Alerting: Update provisioning services that handle Alertmanager configuraiton to access config via storage (#79814 ) * extract get and save operations to a alertmanagerConfigStore. this removes duplicated code in service (currently only mute timings) and improves testing * replace generic errors with errutils one with better messages. * update provisioning services to use new store --------- Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com>	2024-01-05 16:15:18 -05:00
Alexander Weaver	a8fb01a502	Swap weaveworks/common utilities for equivalents in grafana/dskit (#80051 ) * Replace histogram collector and grpc injectors * Extract request timing utility * Also vendor test file * Suppress erroneous linter warn	2024-01-05 10:08:38 -06:00
Matthew Jacobson	3537c5440f	Alerting: Refactor migration to return pairs of legacy and upgraded structs (#79719 ) Some refactoring that will simplify next changes for dry-run PRs. This should be no-op as far as the created ngalert resources and database state, though it does change some logs. The key change here is to modify migrateOrg to return pairs of legacy struct + ngalert struct instead of actually persisting the alerts and alertmanager config. This will allow us to capture error information during dry-run migration. It also moves most persistence-related operations such as title deduplication and folder creation to the right before we persist. This will simplify eventual partial migrations (individual alerts, dashboards, channels, ...). Additionally it changes channel code to deal with PostableGrafanaReceiver instead of PostableApiReceiver (integration instead of contact point).	2024-01-05 05:37:13 -05:00
Santiago	1f6575e65e	Alerting: Test MOA in remote secondary mode (#79828 )	2024-01-05 11:05:27 +01:00
Alexander Weaver	90d4704cd7	Alerting: Fix URL timestamp conversion in historian API in annotation mode (#80026 ) Fix timestamp conversion when calling annotation store	2024-01-04 12:40:21 -06:00
Yuri Tseretyan	f6a46744a6	Alerting: Support hysteresis command expression (#75189 ) Backend: * Update the Grafana Alerting engine to provide feedback to HysteresisCommand. The feedback information is stored in state.Manager as a fingerprint of each state. The fingerprint is persisted to the database. Only fingerprints that belong to Pending and Alerting states are considered as "loaded" and provided back to the command. - add ResultFingerprint to state.State. It's different from other fingerprints we store in the state because it is calculated from the result labels. - add rule_fingerprint column to alert_instance - update alerting evaluator to accept AlertingResultsReader via context, and update scheduler to provide it. - add AlertingResultsFromRuleState that implements the new interface in eval package - update getExprRequest to patch the hysteresis command. * Only one "Recovery Threshold" query is allowed to be used in the alert rule and it must be the Condition. Frontend: * Add hysteresis option to Threshold in UI. It's called "Recovery Threshold" * Add test for getUnloadEvaluatorTypeFromCondition * Hide hysteresis in panel expressions * Refactor isInvalid and add test for it * Remove unnecesary React.memo * Add tests for updateEvaluatorConditions --------- Co-authored-by: Sonia Aguilar <soniaaguilarpeiron@gmail.com>	2024-01-04 11:47:13 -05:00
Santiago	a77ba40ed4	Alerting: Use the forked Alertmanager for remote secondary mode (#79646 ) * (WIP) Alerting: Use the forked Alertmanager for remote secondary mode * fall back to using internal AM in case of error * remove TODOs, clean up .ini file, add orgId as part of remote AM config struct * log warnings and errors, fall back to remoteSecondary, fall back to internal AM only * extract logic to decide remote Alertmanager mode to a separate function, switch on mode * tests * make linter happy * remove func to decide remote Alertmanager mode * refactor factory function and options * add default case to switch statement * remove ineffectual assignment	2023-12-21 15:26:31 +01:00
Santiago	c46da8ea9b	Alerting: Update alerting package and imports from cluster and clusterpb (#79786 ) * Alerting: Update alerting package * update to latest commit * alias for imports	2023-12-21 12:34:48 +01:00
Matthew Jacobson	0424d44b39	Alerting: In migration, create one label per channel (#76527 ) * In migration, create one label per channel This PR changes how routing is done by the legacy alerting migration. Previously, we created a single label on each alert rule that contained an array of contact point names. Ex: __contact__="slack legacy testing","slack legacy testing2" This label was then routed against a series of regex-matching policies with continue=true. Ex: __contacts__ =~ ."slack legacy testing". In the case of many contact points, this array could quickly become difficult to manage and difficult to grok at-a-glance. This PR replaces the single __contact__ label with multiple __legacy_c_{contactname}__ labels and simple equality-matching policies. These channel-specific policies are nested in a single route under the top-level route which matches against __legacy_use_channels__ = true for ease of organization. This should improve the experience for users wanting to keep the default migrated routing strategy but who also want to modify which contact points an alert sends to.	2023-12-19 13:25:13 -05:00
Santiago	9945514baa	Alerting: Validate configuration for the remote Alertmanager struct (#79691 ) * Alerting: Validate configuration for the remote Alertmanager struct * add TenantID to test * add OrgID to config struct in tests	2023-12-19 18:41:48 +01:00
Alexander Weaver	65ecde6eed	Alerting: Don't record annotations for mapped NoData transitions, when NoData is mapped to OK (#77164 ) * Exclude mapped nodata transitions when nodata mapped to OK * Fix processEvalResults test * Don't check NoDataState when filtering transition * Add comment to explain purpose of separate function --------- Co-authored-by: William Wernert <william.wernert@grafana.com>	2023-12-18 16:59:32 -05:00
Santiago	f7248efff5	Alerting: Fix panic when creating a new Alertmanager returns an error (#79641 ) Alerting: Fix panic after error creating new Alertmanager	2023-12-18 15:33:07 +01:00
Alexander Weaver	cf8e8852c3	Alerting: Drop NamespaceID from responses on unstable ngalert API endpoints in favor of NamespaceUID (#79359 ) * Drop from API response * Drop from swagger docs * Drop from integration tests * regenerate public swagger docs * Drop from frontend * Drop asserts for namespaceID field	2023-12-15 11:06:53 -06:00
William Wernert	9171bf92bb	Alerting: Add rule ID and title to alert state history Loki entry (#79481 ) * Add rule ID and title to Loki entry * Combine related tests	2023-12-14 13:06:23 -05:00
Santiago	23b4568597	Alerting: Send configuration and state to the remote Alertmanager on shutdown (#78682 ) * Alerting: Send configuration and state to the remote Alertmanager on shutdown * Alerting: Add a sync interval for ApplyConfig in remote secondary mode * add routine to sync states and configs * pass a cancellable context to syncRoutine(), remove tests for ApplyConfig, cache last config in memory * extract logic to update config and state in the remote Alertmanager * get latest config from the database * avoid using separate goroutine for updating state and config * clean up PR * refactor, comments, tests * update tests * remove canceled context from calls to StopAndWait() * create context with timeout and send config and state to remote Alertmanager * update tests * address code review comments	2023-12-13 22:53:09 +01:00
Julien Duchesne	884e0427e6	ngalert openapi: Add `X-Disable-Provenance` to missing operations (#79278 ) Swagger(ngalert): Add `X-Disable-Provenance` to missing operations I added all functions that call the `determineProvenance` function Schema changes are from: `make` in `pkg/services/ngalert/api/tooling` `make swagger-clean && make openapi3-gen` in root	2023-12-13 10:55:59 -05:00
Santiago	91836e7832	Alerting: Add time-based convergence in remote secondary mode (#78809 ) * Alerting: Add a sync interval for ApplyConfig in remote secondary mode * add routine to sync states and configs * pass a cancellable context to syncRoutine(), remove tests for ApplyConfig, cache last config in memory * extract logic to update config and state in the remote Alertmanager * get latest config from the database * avoid using separate goroutine for updating state and config * clean up PR * refactor, comments, tests * update tests * add config struct for remote secondary forked Alertmanager * use errgroups for sync operations * use waitgroup instead of errgroup * remove helper method to sync AMs * check for errors instead of bool syncErr	2023-12-13 13:36:17 +01:00
William Wernert	62bdbe5b44	Annotations/Alerting: Add Loki historian store stub (#78363 ) * Add Loki historian store stub * Add composite store * Use composite store if Loki historian enabled * Split store interface into read/write * Make composite + historian stores read only * Use variadic constructor for composite * Modify Loki store enable logic * Use dskit.concurrency.ForEachJob for parallelism	2023-12-12 17:43:09 -05:00
Alexander Weaver	aa63e91a43	Alerting: Use mux router to match hooks, add support for path variables and methods (#79345 ) * Use a router inside hooks rather than plain string matching * Add test for mismatched method	2023-12-12 14:43:11 -06:00
Julien Duchesne	f977e3faf5	ngalert swagger: Fix status code (#79415 ) This endpoint returns a 202, not a 204 Let me know if we should instead change the response of the API	2023-12-12 13:40:36 -05:00
Santiago	1a5c2cb55b	Alerting: Check whether the internal Alertmanager is ready in remote secondary mode (#79406 ) Alerting: Check whether the internal Alertmanager is ready in remote secondary	2023-12-12 18:33:11 +01:00
gotjosh	cc3c0a2cc2	Alerting: Refactor readiness check (#78799 ) * Alerting: Refactor readiness check Moves the readiness check to the mimir client and removes the need to assert that we have senders - it already has a queue and can hold notifications until we're ready to send them. --------- Signed-off-by: gotjosh <josue.abreu@gmail.com>	2023-12-12 15:34:54 +00:00
Santiago	57e0d6bcb5	Chore: Simplify function signature for GetLatestAlertmanagerConfiguration (#79392 )	2023-12-12 13:49:54 +01:00
Yuri Tseretyan	8af08d0df2	Alerting: Add export of mute timings to file provisioning formats (#79225 ) * add export of mute timings to file provisioning formats * support export of mute timings to HCL	2023-12-11 21:36:51 -05:00
Alexander Weaver	b867505bd4	alerting: Add tests for hooks (#79284 ) Add tests for hooks	2023-12-11 13:20:48 -06:00
Yuri Tseretyan	2be7605794	Alerting: Fix fine-grained rule access control to use 403 for authorization error (#79239 ) * use 403 for authorization error * update silences API * add ForbiddenError to rule API responses	2023-12-07 13:43:58 -05:00
gotjosh	c631261681	Alerting: Attempt to retry retryable errors (#79161 ) * Alerting: Attempt to retry retryable errors Retrying has been broken for a good while now (at least since version 9.4) - this change attempts to re-introduce them in their simplest and safest form possible. I first introduced #79095 to make sure we don't disrupt or put additional load on our customer's data sources with this change in a patch release. Paired with this change, retries can now work as expected. There's two small differences between how retries work now and how they used to work in legacy alerting. Retries only occur for valid alert definitions - if we suspect that that error comes from a malformed alert definition we skip retrying. We have added a constant backoff of 1s in between retries. --------- Signed-off-by: gotjosh <josue.abreu@gmail.com>	2023-12-06 20:45:08 +00:00
gotjosh	07915703fe	Revert "Alerting: Attempt to retry retryable errors" (#79158 ) Revert "Alerting: Attempt to retry retryable errors (#79037)" This reverts commit `3e51cf0949`.	2023-12-06 19:12:01 +00:00
gotjosh	3e51cf0949	Alerting: Attempt to retry retryable errors (#79037 ) * Alerting: Attempt to retry retryable errors Currently in a draft state, but this was the minimal diff I could put together to exemplify how could achieve this. Signed-off-by: gotjosh <josue.abreu@gmail.com> --------- Signed-off-by: gotjosh <josue.abreu@gmail.com>	2023-12-06 16:35:22 +00:00
Yuri Tseretyan	7e331c8507	Alerting: Support for `condition` field in /api/v1/eval (#79032 ) Co-authored-by: Sonia Aguilar <soniaaguilarpeiron@gmail.com>	2023-12-06 11:28:43 -05:00
Alexander Zobnin	959ebf82da	Folders: Show dashboards and folders with directly assigned permissions in "Shared" folder (#78465 ) * Folders: Show folders user has access to at the root level * Refactor * Refactor * Hide parent folders user has no access to * Skip expensive computation if possible * Fix tests * Fix potential nil access * Fix duplicated folders * Fix linter error * Fix querying folders if no managed permissions set * Update benchmark * Add special shared with me folder and fetch available non-root folders on demand * Fix parents query * Improve db query for folders * Reset benchmark changes * Fix permissions for shared with me folder * Simplify dedup * Add option to include shared folder permission to user's permissions * Fix nil UID * Remove duplicated folders from shared list * Folders: Fix fetching empty folder * Nested folders: Show dashboards with directly assigned permissions * Fix slow dashboards fetch * Refactor * Fix cycle dependencies * Move shared folder to models * Fix shared folder links * Refactor * Use feature flag for permissions * Use feature flag * Review comments * Expose shared folder UID through frontend settings * Add frontend type for sharedWithMeFolderUID option * Refactor: apply review suggestions * Fix parent uid for shared folder * Fix listing shared dashboards for users with access to all folders * Prevent creating folder with "shared" UID * Add tests for shared folders * Add test for shared dashboards * Fix linter * Add metrics for shared with me folder * Add metrics for shared with me dashboards * Fix tests * Tests: add metrics as a dependency * Fix access control metadata for shared with me folder * Use constant for shared with me * Optimize parent folders access check, fetch all folders in one query. * Use labels for metrics	2023-12-05 16:13:31 +01:00
Rodrigo Villablanca	ab83bc7346	Alerting: Fix export of notification policy to JSON (#78021 ) * Export Notification Policy correctly (#78020) The JSON version of an exported Notification Policy now inline correctly the policy in the same way the Yaml version does. Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2023-12-04 16:57:37 -05:00
Julien Duchesne	3c51190392	ngalert `make`: Support GNU install on Darwin (#78482 ) * ngalert `make`: Support GNU install on Darwin Currently, the Makefile assumes that Darwin is using the Mac version of `sed` I have the GNU version, so it failed. With this PR, it checks which version is installed I also called `make` and there are some changes that came out of it * swagger-gen	2023-12-04 10:11:39 -05:00
Sofia Papagiannaki	6d4625ad52	Alerting: Fix deleting rules in a folder with matching UID in another organization (#78258 ) * Remove usage of obsolete function for deleting alert rules under folder * Apply suggestion from code review * Update tests	2023-12-04 11:34:38 +02:00
Yuri Tseretyan	64feeddc23	Alerting: Update rule access control to return errutil errors (#78284 ) * update rule access control to return errutil errors * use alerting in msgID	2023-12-02 01:42:11 +02:00
Alexander Weaver	ab0ef5276f	Alerting: Decouple quota configuration logic from API interfaces and add tests (#78930 ) * Separate usage reporter from API * Extract quota registration * Decouple from API store interface * Move to ngalert package and add tests * linter	2023-12-01 10:47:19 -06:00
Steve Simpson	520c927931	Alerting: Only warm alert state cache if execute_alerts=true. (#78895 ) * Alerting: Only warm alert state cache if execute_alerts=true. If the Grafana instance is not executing alerts, then Warm()-ing the state manager is wasteful and could lead to misleading rule status queries, as the status returned will be always based on the state loaded from the database at startup, and not the most recent evaluation state. * Move Warm() down to shared conditional.	2023-12-01 10:17:32 +01:00
Matthew Jacobson	5a80962de9	Alerting: Add clean_upgrade config and deprecate force_migration (#78324 ) * Alerting: Add clean_upgrade config and deprecate force_migration Upgrading to UA and rolling back will no longer delete any data by default. Instead, each set of tables will remain unchanged when switching between legacy and UA. As such, the force_migration config has been deprecated and no extra configuration is required to roll back to legacy anymore. If clean_upgrade is set to true when upgrading from legacy alerting to Unified Alerting, grafana will first delete all existing Unified Alerting resources, thus re-upgrading all organizations from scratch. If false or unset, organizations that have previously upgraded will not lose their existing Unified Alerting data when switching between legacy and Unified Alerting. Similar to force_migration, it should be kept false when not needed as it may cause unintended data-loss if left enabled. --------- Co-authored-by: Christopher Moyer <35463610+chri2547@users.noreply.github.com>	2023-11-30 11:01:11 -05:00
Matthew Jacobson	cdad712547	Alerting: Keep track of individual org migration status (#78369 ) * Alerting: Keep track of individual org migration status Save migration status per migrated org. Change the meaning (and key/value) of the org_id=0 entry to store the current (previous) config value used by alerting. This is so we can know when to upgrade/downgrade by comparing with the new config value in UnifiedAlerting.IsEnabled.	2023-11-30 10:25:59 -05:00
Santiago	d64c2b6f4e	Alerting: Implement ApplyConfig in the forked Alertmanager (#78684 ) * Alerting: Add a sync interval for ApplyConfig in remote secondary mode * remove out of scope code * remove parentheses after CleanUp for consistency in test comments * Add comment to ApplyConfig	2023-11-30 15:36:41 +01:00
Santiago	316c8b50bc	Alerting: Add SaveAndApply methods to the forked Alertmanager (remote secondary) (#78827 ) * Alerting: Add configuration methods to the forked Alertmanager for remote secondary modes * update comments	2023-11-30 15:18:56 +01:00
Matthew Jacobson	2b51f0e263	Alerting: In migration improve deduplication of title and group (#78351 ) * Alerting: In migration improve deduplication of title and group This change improves alert titles generated in the legacy migration that occur when we need to deduplicate titles. Now when duplicate titles are detected we will first attempt to append a sequential index, falling back to a random uid if none are unique within 10 attempts. This should cause shorter and more easily readable deduplicated titles in most cases. In addition, groups are no longer deduplicated. Instead we set them to a combination of truncated dashboard name and humanized alert frequency. This way, alerts from the same dashboard share a group if they have the same evaluation interval. In the event that truncation causes overlap, it won't be a big issue as all alerts will still be in a group with the correct evaluation interval.	2023-11-29 10:05:00 -05:00
Santiago	73776f37eb	Alerting: Send state to the remote Alertmanager (#78538 ) * Alerting: Introduce a Mimir client as part of the Remote Alertmanager Mimir client that understands the new APIs developed for mimir. Very much a WIP still. * more wip * appease the linter * more linting * add more code * get state from kvstore, encode, send * send state to the remote Alertmanager, extract fullstate logic into its own function * pass kvstore to remote.NewAlertmanager() * refactor * add fake kvstore to tests * tests * use FileStore to get state * always log 'completed state upload' * refactor compareRemoteConfig * base64-encode the state in the file store * export silences and nflog filenames, refactor * log 'completed state/config upload...' regardless of outcome * add values to the state store in tests * address code review comments * log error from filestore --------- Co-authored-by: gotjosh <josue.abreu@gmail.com>	2023-11-29 12:49:39 +01:00
Matthew Jacobson	ce90a1f2be	Alerting: Apply query optimization to eval endpoints (#78566 ) * Alerting: Apply query optimization to eval endpoints Previously, query optimization was applied to alert queries when scheduled but not when ran through `api/v1/eval` or `/api/v1/rule/test/grafana`. This could lead to discrepancies between preview and scheduled alert results.	2023-11-28 19:44:28 -05:00
Santiago	01d274852c	Alerting: Add GetFullState method to FileStore (#78701 ) * Alerting: Add GetFullState method to FileStore * make tests compile, create stateStore in NewAlertmanager * return errors instead of logging, accept an arbitrary number of strings * make NewAlertmanager() accept a stateStore	2023-11-28 15:34:45 +01:00
William Wernert	f7bf818527	Alerting: Make alert state history Loki http client public (#78291 ) * Make state history Loki client public * Make historian metrics subsystem configurable	2023-11-27 09:20:50 -05:00
Matthew Jacobson	4b439b7f52	Alerting: In migration, fallback to '1s' for malformed min interval (#78614 ) * Alerting: In migration, fallback to '1s' for malformed min interval During legacy migration, when we encounter an alert datasource query with a min interval (interval field in the query model) that is not parseable, instead of failing the migration we fallback to a min interval of 1s and continue. The reason for this is a bug in legacy alerting (existing for a few major versions) which allows arbitrary dashboard variables to be used as the min interval, even though those variables do not work and will cause the legacy alert to fail with `interval calculation failed: time: invalid duration`.	2023-11-24 11:27:44 -05:00
gotjosh	8120306fea	Remote Alertmanager(refactor): Only parse the URL once (#78631 ) * Remote Alertmanager(refactor): Only parse the URL once Exactly what it says in the tin. Signed-off-by: gotjosh <josue.abreu@gmail.com> * use the existing tests Signed-off-by: gotjosh <josue.abreu@gmail.com> --------- Signed-off-by: gotjosh <josue.abreu@gmail.com>	2023-11-24 11:05:13 +00:00
Jean-Philippe Quéméner	11d4f604f5	fix(alerting): proper handling for queries with multiple conditions in migration (#78591 ) fix(alerting): proper handling for queries with multiple conditions	2023-11-23 18:05:44 +01:00
gotjosh	23fe8f4e9c	Alerting: Introduce a Mimir client as part of the Remote Alertmanager (#78357 ) * Alerting: Introduce a Mimir client as part of the Remote Alertmanager This is our first attempt at making Grafana communicate use Mimir as a backend - it uses a new set of APIs that we've developed on the Mimir side to upload the grafana configuration and alertmanager state so that it can then be ported over. Codewise, we've introduced a couple of things: A client to isolate in its own package all the communication that happens with Mimir A few changes to the remote/alertmanager to include uploading the configuration and state when it starts A few refactors that align a bit better with the design approach that we're thinking An integration tests again these newly developed APIs using a custom image --------- Signed-off-by: gotjosh <josue.abreu@gmail.com> Co-authored-by: Santiago <santiagohernandez.1997@gmail.com>	2023-11-23 16:59:36 +00:00
Jo	0de66a8099	Authz: Remove use of SignedInUser copy for permission evaluation (#78448 ) * remove use of SignedInUserCopies * add extra safety to not cross assign permissions unwind circular dependency dashboardacl->dashboardaccess fix missing import * correctly set teams for permissions * fix missing inits * nit: check err * exit early for api keys	2023-11-22 14:20:22 +01:00
Tania	39754ba2d6	Nested Folders: Wrap create/update operations with transactions (#78000 ) * Nested Folders: Add transaction to create and update methods * Update tests * Make IncreaseVersionForAllRulesInNamespace synchronous * Resolve merge conflicts	2023-11-21 23:06:20 +02:00
Kat Yang	2f2ce3edbb	Chore: Deprecate ID from Folder (#78281 ) * Chore: Deprecate ID from Folder * chore: add more linter comments * chore: add missing lint comment	2023-11-20 15:44:51 -05:00
Matthew Jacobson	893839d27b	Alerting: Move general alert rule validation from db-layer to model (#78325 ) Alerting: Move general alert rule validation to model	2023-11-17 11:20:50 -05:00
Jean-Philippe Quéméner	2d2e058563	refactor: use constant for prometheus datasource type (#78287 )	2023-11-17 01:07:35 +01:00
Yuri Tseretyan	7cec741bae	Alerting: Extract alerting rules authorization logic to a service (#77006 ) * extract alerting authorization logic to separate package * convert authorization logic to service	2023-11-15 18:54:54 +02:00
Kat Yang	3a2e96b0db	Chore: Deprecate FolderID from Dashboard (#77823 ) * Chore: Deprecate FolderID from Dashboard * chore: add two missing nolint comments	2023-11-15 10:28:50 -05:00
Ryan McKinley	f69fd3726b	FeatureToggles: Add context and and an explicit global check (#78081 )	2023-11-14 12:50:27 -08:00
Jo	580477bf8e	NGAlerting: Use identity.Requester interface instead of SignedInUser (#76360 ) * unfurl SignedInUserAttrs services * replace signedInUser with Requester replace signedInUser with requester * fix tests * linting --------- Co-authored-by: Ieva <ieva.vasiljeva@grafana.com>	2023-11-14 14:47:34 +00:00
Santiago	4a152a0e35	Alerting: Add lifecycle methods to the forked Alertmanager (#77741 ) * Alerting: Add an empty Forked Alertmanager * Alerting: Add methods for silences to the forked Alertmanager * check for errors in tests * make linter happy * Alerting: Add methods for alerts to the forked Alertmanager * Alerting: Add methods for receivers to the forked Alertmanager * Alerting: Add TestTemplate method to the forked Alertmanager * make linter happy * separate into both forked AMs * fix tests * Alerting: Add lifecycle methods to the forked Alertmanager	2023-11-14 11:17:17 +01:00
Ryan McKinley	dec9a07738	Settings: Actually deprecate access to feature flags (#78073 )	2023-11-13 11:39:01 -08:00
Ryan McKinley	3509a5abb9	FeatureFlags: Cleanup usage of cfg.IsFeatureToggleEnabled (#78014 )	2023-11-13 07:55:15 -08:00
Santiago	8b751eb216	Alerting: Add TestTemplate method to the forked Alertmanager (#77577 ) * Alerting: Add an empty Forked Alertmanager * Alerting: Add methods for silences to the forked Alertmanager * check for errors in tests * make linter happy * Alerting: Add methods for alerts to the forked Alertmanager * Alerting: Add methods for receivers to the forked Alertmanager * Alerting: Add TestTemplate method to the forked Alertmanager * make linter happy * separate into both forked AMs * fix tests	2023-11-09 12:35:24 +01:00
Santiago	ba51c371ec	Alerting: Add methods for receivers to the forked Alertmanager (#77574 ) * Alerting: Add an empty Forked Alertmanager * Alerting: Add methods for silences to the forked Alertmanager * check for errors in tests * make linter happy * Alerting: Add methods for alerts to the forked Alertmanager * Alerting: Add methods for receivers to the forked Alertmanager * make linter happy * separate into both forked AMs * fix tests * rename testErr -> expErr	2023-11-09 11:38:16 +01:00
Santiago	e24fe96d90	Alerting: Add methods for alerts to the forked Alertmanager (#77571 ) * Alerting: Add an empty Forked Alertmanager * Alerting: Add methods for silences to the forked Alertmanager * check for errors in tests * make linter happy * Alerting: Add methods for alerts to the forked Alertmanager * make linter happy * separate into both forked AMs * rename testErr -> expErr	2023-11-08 13:52:04 +01:00
Santiago	197f0d2859	Alerting: Add methods for silences to the forked Alertmanager (#77805 ) * Alerting: Add an empty Forked Alertmanager * Alerting: Add methods for silences to the forked Alertmanager * check for errors in tests * make linter happy * make linter happy * Alerting: Add methods for silences to the forked Alertmanager	2023-11-08 12:03:40 +01:00
Yuri Tseretyan	a2629f3dd3	Alerting: Remove unused Accesscontrol dependency from DbStore (#77479 )	2023-11-02 15:54:30 -04:00
William Wernert	e562250f72	Alerting: Handle edge cases without panicking during template migration (#76890 ) * Handle empty variable, remove panics * Use fmt.Errorf only where appropriate	2023-11-02 13:24:54 -04:00
Santiago	01af8f61f1	Alerting: Separate the forked Alertmanager into two implementations (#77582 )	2023-11-02 17:53:18 +01:00
Santiago	8fc9873443	Alerting: Add an empty Forked Alertmanager struct (#77550 ) Alerting: Add an empty Forked Alertmanager	2023-11-02 16:49:03 +01:00
Yuri Tseretyan	85425b2194	Alerting: Fix flaky test TestExportRules (#77519 ) * fix test to correclty mock data store * Update pkg/services/ngalert/api/api_ruler_export_test.go Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com> * Update pkg/services/ngalert/api/api_ruler_export_test.go --------- Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com>	2023-11-01 21:35:04 +02:00
Ryan McKinley	5d5f8dfc52	Chore: Upgrade Go to 1.21.3 (#77304 )	2023-11-01 09:17:38 -07:00
Santiago	a6b9b27673	Alerting: Remove OrgID() from the Alertmanager interface (#77398 )	2023-10-31 10:58:47 +01:00
Kyle Brandt	e4d1fdc3d0	Errors: Make errors the same in dev as prod (#77366 ) When running in dev mode, error messages would contain an additional "error" property alongside "message". Since this causes confusion, that has been removed and now error messages are the same both modes (using "message").	2023-10-30 14:06:26 -04:00
Yuri Tseretyan	48b55f39bf	Alerting: Add support for responders to Opsgenie integration (#77159 ) * add support for responders in opsgenie UI config * update export model Co-authored-by: Santiago <santiagohernandez.1997@gmail.com>	2023-10-27 13:06:46 -04:00
Santiago	f9fc2e4568	Alerting: Remove ConfigHash() from the Alertmanager interface (#77134 )	2023-10-25 17:11:53 +02:00
Alexander Weaver	6ee52ac80c	Alerting: Allow more time before Alertmanager expire-resolves alerts (#77094 ) * Sync endsAt factor with prometheus * Fix state tests	2023-10-25 10:03:46 -05:00
Santiago	322a9c0b15	Alerting: Replace FileStore() for CleanUp() in the Alertmanager interface (#77126 ) Alerting: Remplace FileStore() for CleanUp() in the Alertmanager interface	2023-10-25 13:58:28 +02:00
Santiago	01add144b8	Alerting: Send alerts to the remote Alertmanager (#77034 ) * Alerting: Rename remote.ExternalAlertmanager to remote.Alertmanager * Alerting: Send alerts to the remote Alertmanager * add ticker to readiness check, add tests * use options when creating a new sender.ExternaAlertmanager * unexport defaultMaxQueueCapacity * delete unused defaultConfig field * add debug log line when sending alerts to the remote alertmanager * move and refactor readiness check * update tests to not include defaultConfig	2023-10-25 11:52:48 +02:00
Alexander Weaver	39599fa7f7	Alerting: Alert rule constraint violations return as 400s in provisioning API (#76396 ) Constraint violations become 400s	2023-10-23 10:28:40 -05:00
Santiago	488a60aee6	Alerting: Rename remote.ExternalAlertmanager to remote.Alertmanager (#76956 )	2023-10-23 15:37:14 +02:00
gotjosh	866acbd5ac	Alerting: Move `ExternalAlertmanager` to its own package (#76854 ) * Alerting: Move `ExternalAlertmanager` to its own package We'll avoid import cycles when using components from other packages. In addition to that, I've created an `Options` approach for the multiorg alertmanger to allow us to override how per tenant alertmanagers are created. * switch things around * address review comments * fix references and warnings	2023-10-20 14:08:13 +02:00
Santiago	a60ec150f9	Alerting: Fetch receivers from remote Alertmanager (#76841 ) * Alerting: fetch receivers from remote Alertmanager * make linter happy * change require.Eventually() timeout and tick	2023-10-20 11:34:17 +02:00
Steve Simpson	a0476741f2	Alerting: Fix HCL export for alerts with non-zero "for" field. (#76739 ) * Alerting: Fix HCL export for alerts with non-zero "for" field. Fixes #76734 * fix tests --------- Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2023-10-20 11:09:08 +02:00
Matthew Jacobson	c2efcdde09	Alerting: Fix flaky SQLITE_BUSY when migrating with provisioned dashboards (#76658 ) * Alerting: Move migration from background service run to ngalert init sqlite database write contention between the migration's single transaction and dashboard provisioning's frequent commits was causing the migration to fail with SQLITE_BUSY/SQLITE_BUSY_SNAPSHOT on all retries. This is not a new issue for sqlite+grafana, but the discrepancy between the length of the transactions was causing it to be very consistent. In addition, since a failed migration has implications on the assumed correctness of the alertmanager and alert rule definition state, we cause a server shutdown on error. This can make e2e tests as well as some high-load provisioned sqlite installations flaky on startup. The correct fix for this is better transaction management across various services and is out of scope for this change as we're primarily interested in mitigating the current bout of server failures in e2e tests when using sqlite.	2023-10-19 10:03:00 -04:00
Santiago	61cb26711e	Alerting: Fetch alerts from a remote Alertmanager (#75844 ) * Alerting: post alerts to the remote Alertmanager and fetch them * fix broken tests * Alerting: Add Mimir Backend image to devenv (blocks) * add alerting as code owner for mimir_backend block * Alerting: Use Mimir image to run integration tests for the remote Alertmanager * skip integration test when running all tests * skipping integration test when no Alertmanager URL is provided * fix bad host for mimir_backend * remove basic auth testing until we have an nginx image in our CI * add integration tests for alerts * fix tests * change SendCtx -> Send, add context.Context to Send, fix CI * add reover() for functions from the Prometheus Alertmanager HTTP client that could panic * add TODO to implement PutAlerts in a way that mimicks what Prometheus does * fix log format	2023-10-19 11:27:37 +02:00
Alexander Weaver	acee3efcf9	Alerting: Use common StateReason values for NoData/Error mapped states (#76781 ) Fix hardcoded state reasons	2023-10-18 17:26:41 -05:00
Santiago	7d9b2c73c7	Alerting: Use Mimir image to run integration tests for the remote Alertmanager (#76608 ) * Alerting: Use Mimir image to run integration tests for the remote Alertmanager * skip integration test when running all tests * skipping integration test when no Alertmanager URL is provided * fix bad host for mimir_backend * remove basic auth testing until we have an nginx image in our CI	2023-10-17 12:21:45 +02:00
Jean-Philippe Quéméner	2b8c6d66e1	feat(alerting): add query optimizations for prometheus (#76015 )	2023-10-17 11:41:25 +02:00
Torkel Ödegaard	0d55dad075	DashboardScene: Fixes full page reload of fullscreen view of a repeated panel (#76326 ) * Progress on view panel for repeats * Good enough * Update	2023-10-13 16:03:38 +02:00
Matthew Jacobson	a6d928e50e	Alerting: Prevent cleanup of non-empty folders on migration revert (#76439 ) Prevent cleanup of non-empty folders on revert	2023-10-12 18:40:51 -04:00
Matthew Jacobson	5f48619c9a	Alerting: Handle custom dashboard permissions in migration service (#74504 ) * Fix migration of custom dashboard permissions Dashboard alert permissions were determined by both its dashboard and folder scoped permissions, while UA alert rules only have folder scoped permissions. This means, when migrating an alert, we'll need to decide if the parent folder is a correct location for the newly created alert rule so that users, teams, and org roles have the same access to it as they did in legacy. To do this, we translate both the folder and dashboard resource permissions to two sets of SetResourcePermissionCommands. Each of these encapsulates a mapping of all: OrgRoles -> Viewer/Editor/Admin Teams -> Viewer/Editor/Admin Users -> Viewer/Editor/Admin When the dashboard permissions (including those inherited from the parent folder) differ from the parent folder permissions alone, we need to create a new folder to represent the access-level of the legacy dashboard. Compromises: When determining the SetResourcePermissionCommands we only take into account managed and basic roles. Fixed and custom roles introduce significant complexity and synchronicity hurdles. Instead, we log a warning they had the potential to override the newly created folder permissions. Also, we don't attempt to reconcile datasource permissions that were not necessary in legacy alerting. Users without access to the necessary datasources to edit an alert rule will need to obtain said access separate from the migration.	2023-10-12 18:12:40 -04:00
Yuri Tseretyan	372082d254	Alerting: Export of contact points to HCL (#75849 ) * add compat layer to convert from Export model to "new" API models	2023-10-12 22:33:57 +01:00
Yuri Tseretyan	c4ac4eb41b	Alerting: Export of notification policies to HCL (#76411 )	2023-10-12 12:10:08 -04:00
Matthew Jacobson	82f3127e23	Alerting: Move legacy alert migration from sqlstore migration to service (#72702 )	2023-10-12 13:43:10 +01:00
Alexander Weaver	f6649d7a97	Revert "Alerting: Remove vendored models in migration service" (#76387 ) Revert "Alerting: Remove vendored models in migration service (#74503)" This reverts commit `6a8649d544`.	2023-10-11 14:21:21 -05:00
Matthew Jacobson	6a8649d544	Alerting: Remove vendored models in migration service (#74503 ) This PR replaces the vendored models in the migration with their equivalent ngalert models. It also replaces the raw SQL selects and inserts with service calls. It also fills in some gaps in the testing suite around: - Migration of alert rules: verifying that the actual data model (queries, conditions) are correct 9a7cfa9 - Secure settings migration: verifying that secure fields remain encrypted for all available notifiers and certain fields migrate from plain text to encrypted secure settings correctly e7d3993 Replacing the checks for custom dashboard ACLs will be replaced in a separate targeted PR as it will be complex enough alone.	2023-10-11 17:22:09 +01:00
George Robinson	05e12e787b	Alerting: Add provenance field to /api/v1/provisioning/alert-rules (#76252 ) This commit adds the missing Provenance field to responses for /api/v1/provisioning/alert-rules.	2023-10-11 14:51:20 +01:00
Jo	dcd0c6b11e	Identity: Unfurl OrgID in pkg/services to allow using identity.Requester interface (#76113 ) Unfurl OrgID in pkg/services to allow using identity.Requester interface	2023-10-09 10:40:19 +02:00
Yuri Tseretyan	2497db4bd6	Alerting: Add UID of rules to response that were affected by update group request (#75985 ) * update storage's method InstertRules to return ids of added rules as slice to keep the same order as rules in the argument * schematize response of update rule group endpoint, add created, updated, deleted fields that contain UID of affected rules. * update integration tests to use the new fields	2023-10-07 01:11:24 +03:00
Yuri Tseretyan	0a50ca7231	Alerting: Let users with regular permissions access export endpoints (#76082 ) let users with regular permissions access export endpoints	2023-10-06 14:48:20 -04:00
Jo	41bcb5e07f	Identity: Port folder library to identity.Requester (#76105 ) Port folders to identity.Requester	2023-10-06 15:02:34 +02:00
Yuri Tseretyan	4343c99e2a	Add compat function for notify.GrafanaIntegrationConfig to EmbeddedContactPoint (#75995 ) Co-authored-by: Matthew Jacobson <matthew.jacobson@grafana.com>	2023-10-05 23:13:34 +03:00
Yuri Tseretyan	51499d7763	Alerting: Update alert rule export models to omit default values (#75918 ) * do not include rule uid in response if it's empty * make some fields of export models nillable	2023-10-05 15:16:44 -04:00
Yuri Tseretyan	5be52dfe21	Alerting: Fix store's GetNamespaceByUID (#75976 )	2023-10-04 13:13:31 -04:00
Marcus Efraimsson	e4c1a7a141	Tracing: Standardize on otel tracing (#75528 )	2023-10-03 14:54:20 +02:00
Yuri Tseretyan	027bd9356f	Alerting: Rule Modify Export APIs (#75322 ) * extend RuleStore interface to get namespace by UID * add new export API endpoints * implement request handlers * update authorization and wire handlers to paths * add folder error matchers to errorToResponse * add tests for export methods	2023-10-02 11:47:59 -04:00
gotjosh	e877174501	Alerting: Expose metrics for Alertmanager Alerts - `grafana_alerting_alertmanager_alerts` (#75802 ) * Alerting: Expose metrics for Alertmanager Alerts In Grafana, the alert evaluation and alert delivery are combined. We're always used a metric named `grafana_alerting_alerts` to get a sense of what are the alerts that are currently firing (these come from the evaluation side) and opted to not map the alertmanager alerts metric directly. I think it's important that we make a disction between alerts that happen at evaluation vs alerts that are received for delivery by the internal Alertmanager as we have options to skip the delivery of these alerts to the internal alertmanager altogether.	2023-10-02 16:36:23 +01:00
George Robinson	ed7d29f2b9	Alerting: Migrate old alerting templates to Go templates (#62911 ) * Migrate old alerting templates to use $labels * Fix imports * Add test coverage and separate rewriting to Go templates * Fix lint * Check for additional closing braces * Add logging of invalid message templates * Fix tests * Small fixes * Update comments * Panic on empty token * Use logtest.Fake * Fix lint * Allow for spaces in variable names by not tokenizing spaces * Add template function to deduplicate Labels in a Value map * Fix behavior of mapLookupString * Reference deduplicated labels in migrated message template * Fix behavior of deduplicateLabelsFunc * Don't create variable for parent logger * Add more tests for deduplicateLabelsFunc * Remove unused function * Apply suggestions from code review Co-authored by: Yuri Tseretyan <yuriy.tseretyan@grafana.com> * Give label val merge function better name * Extract template migration and escape literal tokens * Consolidate + simplify template migration --------- Co-authored-by: William Wernert <william.wernert@grafana.com>	2023-10-02 11:25:33 -04:00
Santiago	73be9449d1	Alerting: Manage remote Alertmanager silences (#75452 ) * Alerting: Manage remote Alertmanager silences * fix typo * check errors when encoding json in fake external AM * take path from configured URL, check for nil responses	2023-10-02 07:36:11 -03:00
Carl Bergquist	a39d2ae8ea	instrumentation: change slogroup for alerting handlers to high-slow (#75460 ) instrumentation: change slogroup for alerting handlers to high-fast Signed-off-by: bergquist <carl.bergquist@gmail.com>	2023-09-29 14:56:48 +02:00
Yuri Tseretyan	237ce5ea82	Alerting: Extract methods for fetching rule groups with authorization (#75375 ) * extract methods for fetching rule groups with authorization and refactor the request handlers. * add logging to delete handler	2023-09-26 12:45:22 -04:00
gotjosh	59694fb2be	Alerting: Don't use a separate collection system for metrics (#75296 ) * Alerting: Don't use a separate collection system for metrics The state package had a metric collection system that ran every 15s updating the values of the metrics - there is a common pattern for this in the Prometheus ecosystem called "collectors". I have removed the behaviour of using a time-based interval to "set" the metrics in favour of a set of functions as the "value" that get called at scrape time.	2023-09-25 10:27:30 +01:00
William Wernert	925f12d0ea	Alerting: Add support for `keep_firing_for` field from external rulers (#75163 ) * Add support for `keep_firing_for` in ruler proxy * Don't delete `keep_firing_for` when editing a rule with the field set Co-Authored-By: Sonia Aguilar <33540275+soniaAguilarPeiron@users.noreply.github.com> --------- Co-authored-by: Sonia Aguilar <33540275+soniaAguilarPeiron@users.noreply.github.com>	2023-09-21 16:02:53 -04:00
Steve Simpson	894f420014	Alerting: Pass loggers into SchedulerCfg and ManagerCfg. (#75158 )	2023-09-20 15:07:02 +02:00
Santiago	8c1a3f75f9	Alerting: Add empty remote Alertmanager struct (#74864 ) * Alerting: Add empty remote alertmanager struct * Update pkg/services/ngalert/notifier/external_alertmanager.go Co-authored-by: gotjosh <josue.abreu@gmail.com> --------- Co-authored-by: gotjosh <josue.abreu@gmail.com>	2023-09-14 08:55:01 -03:00
Kyle Brandt	35e488b22b	SSE: Localize/Contain Errors within an Expression (#73163 ) Changes SSE to not always fail all queries when one fails. Now only the query itself, and nodes that depend on it will error. --------- Co-authored-by: Gilles De Mey <gilles.de.mey@gmail.com>	2023-09-13 13:58:16 -04:00
Jean-Philippe Quéméner	f3b6d01306	feat(alerting): enable loki query optimization by default (#74739 )	2023-09-13 13:52:40 +02:00
Nutmos	ad9f0b9e4e	Alerting: Add message options for Telegram contact point (#74635 ) Co-authored-by: Santiago <santiagohernandez.1997@gmail.com>	2023-09-12 10:45:57 -04:00
Yuri Tseretyan	6f785f7269	Alerting: Support for single rule and multi-folder rule export (#74625 )	2023-09-11 13:13:02 -04:00
Yuri Tseretyan	dce492642a	Alerting: Export of alert rules in HCL format (#73166 ) * import hashicopr/hcl/v2 * add hcl package and export to HCL * annotate export structs --------- Co-authored-by: Konrad Lalik <konrad.lalik@grafana.com>	2023-09-11 11:48:23 -04:00
Will Browne	e855efb13d	Plugins: Move store and plugin dto to pluginsintegration (#74655 ) move store and plugin dto	2023-09-11 13:59:24 +02:00
Yuri Tseretyan	99fd7b8141	Alerting: Update provisioning to validate user-defined UID on create (#73793 ) * add ValidateUID to util * provisioning to validate UID on rule creation --------- Co-authored-by: brendamuir <100768211+brendamuir@users.noreply.github.com> Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com>	2023-09-08 15:09:35 -04:00
Yuri Tseretyan	0df3647367	Alerting: extend rules export API to filter by folder and group (#74423 ) update endpoint `GET /api/v1/provisioning/alert-rules/export` to accept query parameters `folderUid` and `group`	2023-09-07 17:34:32 -04:00
Santiago	93b9f9b537	Alerting: Use interfaces for the Alertmanager (#73900 )	2023-09-06 07:59:29 -03:00
Alexander Weaver	5c9aeaef41	Alerting: Do not exit if Redis ping fails when using redis-based Alertmanager clustering (#74144 ) Do not fail redis peer construction if ping fails	2023-09-05 10:43:13 -05:00
Ieva	58efa49933	Chore: remove `IsDisabled` method for access control (#74340 ) remove IsDisabled method for access control, clean up tests	2023-09-05 11:04:39 +01:00
Yuri Tseretyan	baea7a7556	Alerting: Fix provisioning of contact points when contact point is renamed (#74238 ) * add test that demonstrates the bug * fix renaming provisioning contact points when it is the last in the group	2023-09-04 13:30:15 -04:00
Serge Zaitsev	58f6648505	Chore: capitalise messages for alerting (#74335 )	2023-09-04 18:46:34 +02:00
github-actions[bot]	eb93ebe0d0	Alerting: Update Swagger spec (#74300 ) chore: update alerting swagger spec Co-authored-by: rwwiv <rwwiv@users.noreply.github.com>	2023-09-04 16:17:49 +00:00
George Robinson	439270f6cb	Rename Google Hangouts to Google Chat (#74162 ) * Rename Google Hangouts to Google Chat * Fix prettier	2023-08-31 16:09:22 +03:00
Ryan McKinley	025b2f3011	Chore: use any rather than interface{} (#74066 )	2023-08-30 18:46:47 +03:00
linoman	1b8e9b51b2	Replace signed in user for identity.requester (#74048 ) * Make identity.Requester available at Context * Clean pkg/services/guardian/guardian.go * Clean guardian provider and guardian AC * Clean pkg/api/team.go * Clean ctxhandler, datasources, plugin and live * Clean dashboards and guardian * Implement NewUserDisplayDTOFromRequester * Change status code numbers for http constants * Upgrade signature of ngalert services * log parsing errors instead of throwing error	2023-08-30 16:51:18 +02:00
github-actions[bot]	42efd13062	Alerting: Update Swagger spec (#73877 ) chore: update alerting swagger spec Co-authored-by: rwwiv <rwwiv@users.noreply.github.com>	2023-08-30 14:00:13 +00:00
Alexander Weaver	dfba94e052	Alerting: Limit redis pool size to 5 and make configurable (#74057 ) * Limit redis pool size to 5 and expose it in config ini * Coerce negative pool sizes to the default	2023-08-29 14:59:12 -05:00
Carl Bergquist	10a82e30ba	Alerting: add route owner middleware (#73869 ) alerting: add route owner middleware Signed-off-by: bergquist <carl.bergquist@gmail.com>	2023-08-29 12:43:33 +02:00
Jo	a307582212	Revert "Replace signed in user for identity.requester (#73750 )" (#73962 ) This reverts commit `9b9c9e83dc`.	2023-08-28 21:05:59 +02:00
linoman	9b9c9e83dc	Replace signed in user for identity.requester (#73750 ) * Make identity.Requester available at Context * Clean pkg/services/guardian/guardian.go * Clean guardian provider and guardian AC * Clean pkg/api/team.go * Clean ctxhandler, datasources, plugin and live * Question: what to do with the UserDisplayDTO? * Clean dashboards and guardian * Remove identity.Requester from ReqContext * Implement NewUserDisplayDTOFromRequester * Fix tests * Change status code numbers for http constants * Upgrade signature of ngalert services * log parsing errors instead of throwing error * Fix tests and add logs * linting	2023-08-28 12:04:36 -05:00
Torkel Ödegaard	3ee26df41e	PublicDashboards: Variables refactor (#73476 ) Co-authored-by: Juan Cabanas <juan.cabanas@grafana.com> Co-authored-by: Ezequiel Victorero <ezequiel.victorero@grafana.com> Co-authored-by: Ryan McKinley <ryantxu@gmail.com>	2023-08-25 13:56:02 -05:00
George Robinson	bbef000202	Alerting: Add contact point for Grafana OnCall (#73733 ) Add contact point for Grafana OnCall	2023-08-24 10:45:12 +02:00
github-actions[bot]	69267cd28b	Alerting: Update Swagger spec (#72568 ) chore: update alerting swagger spec Co-authored-by: rwwiv <rwwiv@users.noreply.github.com>	2023-08-22 14:35:48 -04:00
Misi	d7166f5f96	RBAC: Remove unused scope from alert.instances:read fixed role (#73268 ) Fix alert.instances:read scope for fixed role	2023-08-16 09:55:49 +02:00
Yuri Tseretyan	938e26b59f	Alerting: Add new metrics and tracings to state manager and scheduler (#71398 ) * add metrics and tracing to state manager * propagate tracer to state manager * add scheduler metrics * fix backtesting * add test for state metrics * remove StateUpdateCount * update docs * metrics can be null * add tracer to new tests	2023-08-16 09:04:18 +02:00
Yuri Tseretyan	90e3f516ff	Alerting: Update Discord settings to treat 'url' as a secure setting (#69588 ) * make discord url secure * support migrating unsecure settings to secure settings * Update public/app/features/alerting/unified/utils/receiver-form.ts Co-authored-by: William Wernert <william.wernert@grafana.com> --------- Co-authored-by: Gilles De Mey <gilles.de.mey@gmail.com> Co-authored-by: William Wernert <william.wernert@grafana.com>	2023-08-16 09:03:56 +02:00
Yuri Tseretyan	0717ec11d6	Alerting: Update state manager to change all current states in the case when Error\NoData is executed as Ok\Nomal (#68142 )	2023-08-15 10:27:15 -04:00
Jean-Philippe Quéméner	2266e09f94	Alerting: optimize rules with multiple loki range queries (#73103 )	2023-08-09 19:00:51 +02:00
Yuri Tseretyan	69c8200fc9	Alerting: Add more tests for state manager ProcessEvalResults (#73019 ) Co-authored-by: Matthew Jacobson <matthew.jacobson@grafana.com>	2023-08-09 12:21:12 -04:00
Jo	97ba611e4c	Chore: Fix ngalert Evaluate signature change (#73084 ) fix ngalert Evaluate sig change	2023-08-09 11:27:14 +02:00
Yuri Tseretyan	6b4a9d73d7	Alerting: Export contact points to check access control action instead legacy role (#71990 ) * introduce a new action "alert.provisioning.secrets:read" and role "fixed:alerting.provisioning.secrets:reader" * update alerting API authorization layer to let the user read provisioning with the new action * let new action use decrypt flag * add action and role to docs	2023-08-08 19:29:34 +03:00
Jean-Philippe Quéméner	2c6cf66741	Alerting: Optimize external Loki queries (#73014 )	2023-08-08 15:13:41 +02:00
Yuri Tseretyan	0053b07885	Alerting: Refactor of state manager tests (#72849 ) * calculate cacheID instead of literals * use mocked clocks * advance clocks with the eval results * use clearer timestamp aliases * make expected state labels be more clear to read Co-authored-by: Matthew Jacobson <matthew.jacobson@grafana.com>	2023-08-04 13:39:49 -04:00
Serge Zaitsev	7767ab6f43	Chore: Add folder data migration, fix unique index (#72602 ) * add folder data migration, fix unique index * fix unique index * pass a fake store in tests * pass store into other providers in tests * and now with alerting!	2023-08-01 09:36:37 +02:00
Yuri Tseretyan	c7598cc6fb	Alerting: Add ability to control scheduler tick interval via config (#71980 ) * add ability to control scheduler interval via config * add feature flag `configurableSchedulerTick`	2023-07-26 12:44:12 -04:00
Yuri Tseretyan	5ba164d92b	Alerting: Exclude expression refIDs from NoData state (#72219 )	2023-07-26 11:42:04 -04:00
Yuri Tseretyan	78fc3bcdf4	Alerting: Fix state manager to not keep datasource_uid and ref_id labels in state after Error (#72216 )	2023-07-26 11:41:46 -04:00
Matthew Jacobson	d31d175109	Alerting: Fix contact point testing with secure settings (#72235 ) * Alerting: Fix contact point testing with secure settings Fixes double encryption of secure settings during contact point testing and removes code duplication that helped cause the drift between alertmanager and test endpoint. Also adds integration tests to cover the regression. Note: provisioningStore is created to remove cycle and the unnecessary dependency.	2023-07-25 10:04:27 -04:00
Arati R	20ffbbc41e	NestedFolders: Add library panels counting and deletion to folder registry (#69149 ) * Expose library element service's folder service * Register library panels, add count implementation * Expand folder counts test * Update registry deletion method interface * Allow getting library elements from any folder * Add test for library panel deletion * Add test for library panel counting	2023-07-25 13:05:53 +02:00
github-actions[bot]	24872370b5	Alerting: Update Swagger spec (#72177 ) chore: update alerting swagger spec Co-authored-by: rwwiv <rwwiv@users.noreply.github.com>	2023-07-25 11:34:00 +02:00
Alexander Weaver	8c8b3ecb5b	Alerting: Add dashboardUID and panelID query parameters for loki state history (#72119 ) * read query parameters * Generate loki query from params	2023-07-24 23:46:46 -05:00
Matthew Jacobson	cfb1656968	Alerting: Add notification policy provisioning file export (#70009 ) * Alerting: Add notification policy provisioning file export - Add provisioning API endpoint for exporting notification policies. - Add option in notification policy view ellipsis dropdown for exporting. - Update various provisioning documentation.	2023-07-24 17:56:53 -04:00
Kyle Brandt	1df4d332c9	SSE: Use errutil to show better error messages in prod (#71658 ) - include public message - propagate data source query errors so they are shown as well to which fixes #70026	2023-07-21 06:38:29 -04:00
Alexander Weaver	ff48a145cc	Alerting: Add exported getters for PanelKey fields (#72064 ) Add getters	2023-07-20 15:47:20 -05:00
Yuri Tseretyan	cbbbe2e6f6	SSE: DSNode to update result with names to make each value identifiable by labels (only Graphite and TestData) (#71246 ) * introduce a function checkIfSeriesNeedToBeFixed to scan all value fields in the response and provide a function that updates Series so they can be uniquely identifiable. Only Graphite and TestData are checked. * update `convertDataFramesToResults` to run this function and provide it to WideToMany * update WideToMany to run the fix function if it is not nil	2023-07-20 14:44:12 -04:00
Matthew Jacobson	13121d3234	Alerting: Add contact point provisioning file export (#71692 ) * Add contact point provisioning file export apis * Regenerate api * docs * frontend * add mock to tests * Fix missing row-level export button on viewer role w/ prov. read * Address review comments --------- Co-authored-by: Gilles De Mey <gilles.de.mey@gmail.com>	2023-07-20 14:35:56 -04:00
George Robinson	8dd3eb856d	Alerting: Improve performance of matching captures (#71828 ) This commit updates eval.go to improve the performance of matching captures in the general case. In some cases we have reduced the runtime of the function from 10s of minutes to a couple 100ms. In the case where no capture matches the exact labels, we revert to the current subset/superset match, but with a reduced search space due to grouping captures.	2023-07-20 09:07:00 +01:00
George Robinson	f1af0502db	Alerting: Add tests for matching captures (#71928 ) This commit adds tests for matching captures, which we do not have at present.	2023-07-19 12:52:26 +01:00
George Robinson	89dcaaf049	Alerting: Sort NumberCaptureValues in EvaluationString (#71927 ) This commit changes extractEvalString to sort NumberCaptureValues in ascending order of Var before building the output string. This means that users will see EvaluationString in a consistent order, but also make it possible to assert its output in tests.	2023-07-19 12:09:21 +01:00
Alexander Weaver	d6db9a5b3c	Alerting: Add exported constructor for panelKey (#71872 ) Exported constructor for panelKey	2023-07-18 13:37:43 -05:00
Alexander Weaver	18b910e654	Alerting: Refactor annotation historian to isolate dashboard service dependency (#71689 ) * Refactor annotation historian to isolate dashboard service dependency * Export PanelKey * Don't export parsePanelKey * Remove commented out code	2023-07-18 08:18:55 -05:00
Will Browne	a8577c21ba	Plugins: Migrate PluginStore mock to pre-existing fakes package (#71664 ) * migrate to existing fakes package * fix imports	2023-07-17 10:21:44 +00:00
Yuri Tseretyan	541bfe636d	SSE: Support for ML query node (#69963 ) * introduce a new node-type ML and implement a command outlier that uses ML plugin as a source of data. * add feature flag mlExpressions that guards the feature	2023-07-13 20:37:50 +03:00
Yuri Tseretyan	64aa5465ac	Alerting: do not expand template for labels\annotations if value is not a template (#71492 )	2023-07-12 14:53:40 -04:00
Kyle Brandt	f6a28cadbc	Alerting: (Chore/Instrumentation) Add traceID to logs with contextual logger (#71289 ) Alerting: (Chore) Add traceID to logs with contextual logger	2023-07-11 10:59:52 +02:00
Matthew Jacobson	e3787de470	Alerting: Fix Alertmanager change detection for receivers with secure settings (#71307 ) * Alerting: Make ApplyAlertmanagerConfiguration only decrypt/encrypt new/changed secure settings Previously, ApplyAlertmanagerConfiguration would decrypt and re-encrypt all secure settings. However, this caused re-encrypted secure settings to be included in the raw configuration when applied to the embedded alertmanager, resulting in changes to the hash. Consequently, even if no actual modifications were made, saving any alertmanager configuration triggered an apply/restart and created a new historical entry in the database. To address the issue, this modifies ApplyAlertmanagerConfiguration, which is called by POST `api/alertmanager/grafana/config/api/v1/alerts`, to decrypt and re-encrypt only new and updated secure settings. Unchanged secure settings are loaded directly from the database without alteration. We determine whether secure settings have changed based on the following (already in-use) assumption: Only new or updated secure settings are provided via the POST `api/alertmanager/grafana/config/api/v1/alerts` request, while existing unchanged settings are omitted. * Ensure saving a grafana-managed contact point will only send new/changed secure settings Previously, when saving a grafana-managed contact point, empty string values were transmitted for all unset secure settings. This led to potential backend issues, as it assumed that only newly added or updated secure settings would be provided. To address this, we now exclude empty ('', null, undefined) secure settings, unless there was a pre-existing entry in secureFields for that specific setting. In essence, this means we only transmit an empty secure setting if a previously configured value was cleared. * Fix linting * refactor omitEmptyUnlessExisting * fixup --------- Co-authored-by: Gilles De Mey <gilles.de.mey@gmail.com>	2023-07-11 08:23:07 +02:00
Yuri Tseretyan	30fc075cd7	Alerting: Fix panic in backtesting API when the testing interval is not times of evaluation interval (#68727 ) * add test for the bug * update backtesting evaluators to accept a number of evaluations instead of `to` to have control over the number evaluations in one place	2023-07-06 11:21:03 -04:00
Yuri Tseretyan	ada325de2a	Alerting: Use unsafe.Slice for hashing a string during rule fingerprint calculation (#71000 )	2023-06-30 14:58:23 -04:00
Alexander Weaver	f94fb765b5	Alerting: Add limit query parameter to Loki-based ASH api, drop default limit from 5000 to 1000, extend visible time range for new ASH UI (#70769 ) * Add limit query parameter * Drop copy paste comment * Extend history query limit to 30 days and 250 entries * Fix history log entries ordering * Update no history message, add empty history test --------- Co-authored-by: Konrad Lalik <konrad.lalik@grafana.com>	2023-06-28 13:32:28 -05:00
George Robinson	594c851d4b	Alerting: Add duration to saving alert states done (#70844 )	2023-06-28 15:19:21 +01:00
Steve Simpson	21ac224c45	Alerting: Make ImageService public in NGAlert. (#70737 )	2023-06-27 13:11:22 +02:00
João Calisto	1d68f5ba77	Alerting: Fix HA alerting membership sync (#70607 ) * Alerting: Fix HA alerting membership sync * Added comment about filtering duplicates	2023-06-26 17:12:10 +01:00
William Wernert	4aa477f48f	Alerting: Move rule UID from Loki stream labels into log lines (#70637 ) Move rule uid into log line to reduce cardinality	2023-06-26 09:57:45 -04:00
George Robinson	7edbe72483	Alerting: Support concurrent queries for saving alert instances (#70525 ) This commit adds support for concurrent queries when saving alert instances to the database. This is an experimental feature in response to some customers experiencing delays between rule evaluation and sending alerts to Alertmanager, resulting in flapping. It is disabled by default.	2023-06-23 11:36:07 +01:00
guangwu	bbe4b0d3de	chore: remove refs to deprecated io/ioutil (#70300 )	2023-06-22 12:19:23 +02:00
Andreas Deininger	95b1f3c875	Fixing typos (#70487 )	2023-06-22 09:43:38 +01:00
Santiago	d3bb9fbbaf	Alerting: Use only token for images in notifications (#70196 ) * Alerting: Use only tokens for images in notifications * update tests * make linter and modfile validator happy	2023-06-21 20:53:45 -03:00
Santiago	ff9eff49bd	Alerting: Bump grafana/alerting and refactor the ImageStore/Provider to provide image URL/bytes (#70182 ) * implement alerting.images.Provider interface in our ImageStore * add URLExists() method to fakeConfigStore * make linter happy * update integration tests	2023-06-21 20:53:30 -03:00
Alexander Weaver	ce6f73bd32	Alerting: Add two missing tests which cover missing URLs for Loki state history (#70460 ) Add two missing tests which cover individual missing URLs	2023-06-21 12:58:37 -05:00

... 2 3 4 5 6 ...

1402 Commits