grafana

mirror of https://github.com/grafana/grafana.git synced 2025-02-25 18:55:37 -06:00

Author	SHA1	Message	Date
Santiago	c9bb18101c	Alerting: Decrypt secrets before sending configuration to the remote Alertmanager (#83640 ) * (WIP) Alerting: Decrypt secrets before sending configuration to the remote Alertmanager * refactor, fix tests * test decrypting secrets * tidy up * test SendConfiguration, quote keys, refactor tests * make linter happy * decrypt configuration before comparing * copy configuration struct before decrypting * reduce diff in TestCompareAndSendConfiguration * clean up remote/alertmanager.go * make linter happy * avoid serializing into JSON to copy struct * codeowners	2024-03-19 12:12:03 +01:00
Matthew Jacobson	3ea5c08c88	Alerting: External AM fix parsing basic auth with escape characters (#84681 )	2024-03-18 13:04:57 -04:00
William Wernert	97f37b2e6f	Alerting: Clamp Loki ASH range query to configured max_query_length (#83986 ) * Clamp range in loki http client to configured max_query_length Defaults to 721h to match Loki default	2024-03-15 18:59:45 +02:00
Yuri Tseretyan	827860d459	Alerting: Alerting accesscontrol utilities (#84508 ) * create fake for accesscontrol.RuleService * make errAuthorizationGeneric public	2024-03-14 14:03:53 -04:00
Yuri Tseretyan	f7d836feed	Alerting: Update rule provisioning service to accept user (#84480 )	2024-03-14 12:04:10 -04:00
Gilles De Mey	8765c48389	Alerting: Remove legacy alerting (#83671 ) Removes legacy alerting, so long and thanks for all the fish! 🐟 --------- Co-authored-by: Matthew Jacobson <matthew.jacobson@grafana.com> Co-authored-by: Sonia Aguilar <soniaAguilarPeiron@users.noreply.github.com> Co-authored-by: Armand Grillet <armandgrillet@users.noreply.github.com> Co-authored-by: William Wernert <rwwiv@users.noreply.github.com> Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-03-14 15:36:35 +01:00
William Wernert	8690a42e33	Alerting: Disallow invalid rule namespace UIDs in provisioning API (#83938 ) * Disallow invalid rule namespace UIDs in provisioning Reject requests with rules that reference a nonexistent folder or have an empty folder uid	2024-03-14 09:58:25 -04:00
Yuri Tseretyan	cfc3957894	Alerting: move store.ErrAlertRuleGroupNotFound to models package (#84308 ) move ErrAlertRuleGroupNotFound to models to avoid future circular dependencies	2024-03-12 15:38:21 -04:00
Andres Martinez Gotor	265200799d	Chore: Update grafana-plugin-sdk (#84289 )	2024-03-12 17:13:23 +01:00
William Wernert	10dc6c6d75	Alerting: Add "Keep Last State" backend functionality (#83940 ) * Implement keep last state for state transitions * Respect For duration when keeping state * Only keep transition from recording an annotation * Add keep last state option for nodata/error in UI	2024-03-12 10:00:43 -04:00
Alexander Weaver	6c5e94095d	Alerting: Scheduler and registry handle rules by an interface (#84044 ) * export Evaluation * Export Evaluation * Export RuleVersionAndPauseStatus * export Eval, create interface * Export update and add to interface * Export Stop and Run and add to interface * Registry and scheduler use rule by interface and not concrete type * Update factory to use interface, update tests to work over public API rather than writing to channels directly * Rename map in registry * Rename getOrCreateInfo to not reference a specific implementation * Genericize alertRuleInfoRegistry into ruleRegistry * Rename alertRuleInfo to alertRule * Comments on interface * Update pkg/services/ngalert/schedule/schedule.go Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com> --------- Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com>	2024-03-11 22:57:38 +02:00
Alexander Weaver	201f5d3ac9	Alerting: Extract large closures in ruleRoutine (#84035 ) * extract notify * extract resetState * move evaluate metrics inside evaluate * split out evaluate	2024-03-06 16:39:23 -06:00
Alexander Weaver	7a171fd14a	Regenerate openapidocs at 1.21.8 to match ci (#84037 ) * Regenerate openapidocs at 1.21.8 to match ci * Adjust trigger to work on the actual outputted files * Also put go.mod and go.sum in the triggers * manually fix * Make an arbitrary change rather than touching the trigger to force a run * Drop all triggers - run all the time * Print diff - taken from @papagian's PR * Manual fixes to swagger doc --------- Co-authored-by: Ryan McKinley <ryantxu@gmail.com>	2024-03-06 16:08:45 -06:00
gotjosh	948c8c45d6	Alerting: Use Alertmanager types extracted into grafana/alerting (#83824 ) * Alerting: Use Alertmanager types extracted into grafana/alerting We're in the process of exporting all Alertmanager types into grafana/alerting so that they can be imported in the Mimir Alertmanager, without a neeed to import Grafana directly. This change introduces type aliasing for all Alertmanager types based on their 1:1 copy that now live in grafana/alerting. Signed-off-by: gotjosh <josue.abreu@gmail.com> --------- Signed-off-by: gotjosh <josue.abreu@gmail.com>	2024-03-06 20:48:32 +00:00
Alexander Weaver	d5fda06147	Alerting: Decouple rule routine from scheduler (#84018 ) * create rule factory for more complicated dep injection into rules * Rules get direct access to metrics, logs, traces utilities, use factory in tests * Use clock internal to rule * Use sender, statemanager, evalfactory directly * evalApplied and stopApplied * use schedulableAlertRules behind interface * loaded metrics reader * 3 relevant config options * Drop unused scheduler parameter * Rename ruleRoutine to run * Update READMED * Handle long parameter lists * remove dead branch	2024-03-06 13:44:53 -06:00
Alexander Weaver	1bb38e8f95	Alerting: Move ruleRoutine to be a method on ruleInfo (#83866 ) * Move ruleRoutine to ruleInfo file * Move tests as well * swap ruleInfo and scheduler parameters on ruleRoutine * Fix linter complaint, receiver name	2024-03-04 17:15:55 -06:00
Alexander Weaver	f2a9d0a89d	Alerting: Refactor ruleRoutine to take an entire ruleInfo instance (#83858 ) * Make stop a real method * ruleRoutine takes a ruleInfo reference directly rather than pieces of it * Fix whitespace	2024-03-04 15:15:01 -06:00
Matthew Jacobson	2e8c514cfd	Alerting: Stop persisting user-defined templates to disk (#83456 ) Updates Grafana Alertmanager to work with new interface from grafana/alerting#161. This change stops passing user-defined templates to the Grafana Alertmanager by persisting them to disk and instead passes them by string.	2024-03-04 20:12:49 +02:00
Alexander Weaver	fa51724bc6	Alerting: Move alertRuleInfo and tests to new files (#83854 ) Move ruleinfo and tests to new files	2024-03-04 11:24:49 -06:00
Ryan McKinley	3036b50df3	Expressions: expose ConvertDataFramesToResults (#83805 )	2024-03-04 18:22:56 +02:00
Ryan McKinley	5f6bf93dd5	Expressions: Use enumerations rather than strings (#83741 )	2024-03-01 19:38:32 +02:00
Santiago	8ad367e4ad	Chore: Remove redundant error check (#83769 )	2024-03-01 13:28:08 -03:00
Alexander Weaver	a862a4264d	Alerting: Export rule validation logic and make it portable (#83555 ) * ValidateInterval doesn't need the entire config * Validation no longer depends on entire folder now that we've dropped foldertitle from api * Don't depend on entire config struct * Export validate group	2024-02-28 14:40:13 -06:00
Joe Blubaugh	b905777ba9	Alerting: Support deleting rule groups in the provisioning API (#83514 ) * Alerting: feat: support deleting rule groups in the provisioning API Adds support for DELETE to the provisioning API's alert rule groups route, which allows deleting the rule group with a single API call. Previously, groups were deleted by deleting rules one-by-one. Fixes #81860 This change doesn't add any new paths to the API, only new methods. --------- Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-02-28 10:19:02 -05:00
김은빈	96dfb385ca	Grafana: Replace magic number with a constant variable in response status (#80132 ) * Chore: Replace response status with const var * Apply suggestions from code review Co-authored-by: Sofia Papagiannaki <1632407+papagian@users.noreply.github.com> * Add net/http import --------- Co-authored-by: Sofia Papagiannaki <1632407+papagian@users.noreply.github.com>	2024-02-27 18:39:51 +02:00
Takashi Idobe	f209a5a8b8	fix typos (#83414 )	2024-02-26 10:52:44 -07:00
Serge Zaitsev	d0679f0993	Chore: Add support bundle for folders (#83360 ) * add support bundle for folders * fix ProvideService in tests * add a test for collector	2024-02-26 11:27:22 +01:00
George Robinson	a0353b237a	Alerting: Update swagger specs (#83260 )	2024-02-23 11:25:43 +00:00
George Robinson	a564c8c439	Alerting: Keep order of time and mute time intervals consistent (#83257 )	2024-02-22 16:57:20 +00:00
George Robinson	1ed1242358	Alerting: Basic support for time_intervals (#83216 ) This commit adds basic support for time_intervals, as mute_time_intervals is deprecated in Alertmanager and scheduled to be removed before 1.0. It does not add support for time_intervals in API or file provisioning, nor does it support exporting time intervals. This will be added in later commits to keep the changes as simple as possible.	2024-02-22 15:58:56 +00:00
Matthew Jacobson	87ab98ea95	Alerting: Fix panic in provisioning filter contacts by unknown name (#83070 )	2024-02-19 17:30:13 +02:00
Matthew Jacobson	46a77c0074	Alerting: Validate upgraded receivers early to display in preview (#82956 ) Previously receivers were only validated before saving the alertmanager configuration. This is a suboptimal experience for those upgrading with preview as the failed channel upgrade will return an API error instead of being summarized in the table.	2024-02-16 15:17:07 -05:00
William Wernert	fabaff9a24	Alerting: Create metric for rules using simple notifications (#82904 ) --------- Co-authored-by: Matthew Jacobson <matthew.jacobson@grafana.com>	2024-02-16 19:01:49 +02:00
Matthew Jacobson	dfaf6d1e2e	Alerting: Dry-run legacy upgrade on startup (#82835 ) Adds a feature flag (alertingUpgradeDryrunOnStart) that will dry-run the legacy alert upgrade on startup. It is enabled by default. When on legacy alerting, this feature flag will log the results of the legacy alerting upgrade on startup and draw attention to anything in the current legacy alerting configuration that will cause issues when the upgrade is eventually performed. It acts as a log warning for those where action is required before upgrading to Grafana v11 where legacy alerting will be removed.	2024-02-16 11:29:54 -05:00
Matthew Jacobson	e7c6e9c5c9	Alerting: Fix migration edge-case race condition for silences (#81206 ) If the db already has an entry in the kvstore for the silences of an alertmanager before the migration has taken place, then it's possible that the active alertmanager will overwrite the silence file created by the migration before it has a chance to load it into memory. This should not happen normally but is possible in edge-cases. This change opts to bypass the unnecessary step of writing the silences to disk during the migration and instead write them directly to the kvstore. This avoids the race condition entirely and is more correct as we treat the database as the source of truth for AM state.	2024-02-16 10:47:34 -05:00
Gabriel MABILLE	846eadff63	RBAC Search: Replace `userLogin` filter by `namespacedID` filter (#81810 ) * Add namespace ID * Refactor and add tests * Rename maxOneOption -> atMostOneOption * Add ToDo * Remove UserLogin & UserID for NamespaceID Co-authored-by: jguer <joao.guerreiro@grafana.com> * Remove unecessary import of the userSvc * Update pkg/services/accesscontrol/acimpl/service.go * fix 1 -> userID * Update pkg/services/accesscontrol/accesscontrol.go --------- Co-authored-by: jguer <joao.guerreiro@grafana.com>	2024-02-16 11:42:36 +01:00
Sven Kirschbaum	86c618a6d6	Alerting: Escape namespace and group path parameters (#80504 ) Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com>	2024-02-16 09:43:47 +01:00
Matthew Jacobson	118e4a50b7	Alerting: Remove start page of legacy upgrade preview (#82010 ) Alerting: Remove start page of upgrade preview Alerting upgrade page will now always show the summary table even before upgrading any alerts or notification channels. There a few reasons for this: - The information on the start page is redundant as it's now contained in the documentation. - Previously, if some unexpected issue prevented performing a full upgrade, a user would have limited to no means to using the preview tool to help fix the problem. This is because you could not see the summary table until the full upgrade was performed at least once. Now, you can upgrade individual alerts and notification channels from the beginning.	2024-02-15 17:34:00 -05:00
Julien Duchesne	ba63e62311	Alerting: Return provenance of notification templates (#82274 )	2024-02-15 14:35:54 -05:00
William Wernert	b7bbc5058f	Alerting: Don't validate rules on group update if they've only been reordered (#81841 ) --------- Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-02-15 12:03:28 -05:00
Yuri Tseretyan	1eebd2a4de	Alerting: Support for simplified notification settings in rule API (#81011 ) * Add notification settings to storage\domain and API models. Settings are a slice to workaround XORM mapping * Support validation of notification settings when rules are updated * Implement route generator for Alertmanager configuration. That fetches all notification settings. * Update multi-tenant Alertmanager to run the generator before applying the configuration. * Add notification settings labels to state calculation * update the Multi-tenant Alertmanager to provide validation for notification settings * update GET API so only admins can see auto-gen	2024-02-15 09:45:10 -05:00
Alexander Weaver	d4ae10ecc6	Alerting: Small refactor, move unrelated functions out of fetcher (#82459 ) Move unrelated functions out of fetcher	2024-02-14 20:01:32 +02:00
Diego Augusto Molina	ff08c0a790	Chore: improve test readability in ngalert/schedule (#82453 ) Chore: improve test readability	2024-02-14 14:53:32 -03:00
Diego Augusto Molina	9c29e1a783	Alerting: Fix data races and improve testing (#81994 ) * Alerting: fix race condition in (ngalert/sender.ExternalAlertmanager).Run Chore: Fix data races when accessing members of ngalert/state.FakeInstanceStore Chore: Fix data races in tests in ngalert/schedule and enable some parallel tests * Chore: fix linters * Chore: add TODO comment to remove loopvar once we move to Go 1.22	2024-02-14 12:45:39 -03:00
Alexander Weaver	ccb4533a86	Alerting: Remove unused AlertRuleVersionWithPauseStatus (#82383 ) Remove unused AlertRuleVersionWithPauseStatus	2024-02-13 10:56:24 -06:00
Alexander Weaver	99fa064576	Alerting: Emit warning when creating or updating unusually large groups (#82279 ) * Add config for limit of rules per rule group * Warn when editing big groups through normal API * Warn on prov api writes for groups * Wire up comp root, tests * Also add warning to state manager warm * Drop unnecessary conversion	2024-02-13 08:29:03 -06:00
Ryan McKinley	0c6e409350	Chore: Update arrow and prometheus dependencies (#82215 ) * update arrow and prometheus * keep codeowner * use compare * use grafana-plugin-sdk-go v0.210.0 --------- Co-authored-by: ismail simsek <ismailsimsek09@gmail.com>	2024-02-13 01:50:25 +01:00
Karl Persson	1315c67c8b	Team/User: UID migrations (#82298 ) * Add user uid migration to run on every startup to protect against empty values in a upgrade downgrade scenario * Add team uid migration to run on every startup to protect against empty values in a upgrade downgrade scenario * Run team uid migration	2024-02-12 14:48:29 +01:00
Alexander Weaver	5bbe9c6e61	Alerting: Enable group-level rule evaluation jittering by default, remove feature toggle (#82212 ) * remove jitter feature flag * Add an out so users can manually disable jitter * Pass in cfg * Add TODO to remove knob in future	2024-02-09 15:53:58 -06:00
Dan Cech	790e1feb93	Chore: Update test database initialization (#81673 ) * streamline initialization of test databases, support on-disk sqlite test db * clean up test databases * introduce testsuite helper * use testsuite everywhere we use a test db * update documentation * improve error handling * disable entity integration test until we can figure out locking error	2024-02-09 09:35:39 -05:00
Jean-Philippe Quéméner	4dc1ebbb66	fix(alerting): add a proper compare func for location in mute timings (#82153 )	2024-02-08 13:36:09 +01:00
George Robinson	90a26e18db	Alerting: Update Alertmanager to e82436c (#82145 ) This commit updates Alertmanager to commit e82436c, which is based on commit f69a508 from Prometheus Alertmanager.	2024-02-08 11:25:27 +00:00
Matthew Jacobson	dd0ca1263b	Alerting: Include rule uid, title, namespace in unique constraint errors (#82011 ) * Alerting: Include rule_uid, title, namespace_uid in unique constraint errors	2024-02-07 12:55:48 -05:00
Alexander Weaver	843c477899	Alerting: Add exported API to scheduler to access currently loaded rules (#82031 ) * Add exported API to fetch rule definitions from scheduler * Add comment	2024-02-07 09:31:22 -06:00
Yuri Tseretyan	47546a4c72	Alerting: Update API to use folders' full paths (#81214 ) * update GetUserVisibleNamespaces to use FolderSeriver * update GetNamespaceByUID to use FolderService.GetFolders * update GetAlertRulesForScheduling to use FolderService.GetFolders * Update API and GetAlertRulesForScheduling to use the folder's full path * get full path of folder in RouteTestGrafanaRuleConfig * fix escaping of titles for MySQL	2024-02-06 17:12:13 -05:00
Gokhan	cf601fab09	Alerting: Enable sending notifications to a specific topic on Telegram (#79546 ) Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-02-06 17:19:22 +02:00
William Wernert	2ea82af6e7	Alerting: Pass in receiver service to API struct (#81978 )	2024-02-06 16:49:47 +02:00
George Robinson	c8ccc4649c	Alerting: Support UTF-8 (#81512 ) This pull request updates our fork of Alertmanager to commit 65bdab0, which is based on commit 5658f8c in Prometheus Alertmanager. It applies the changes from grafana/alerting#155 which removes the overrides for validation of alerts, labels and silences that we had put in place to allow alerts and silences to work for non-Prometheus datasources. However, as this is now supported in Alertmanager with the UTF-8 work, we can use the new upstream functions and remove these overrides. The compat package is a package in Alertmanager that takes care of backwards compatibility when parsing matchers, validating alerts, labels and silences. It has three modes: classic mode, UTF-8 strict mode, fallback mode. These modes are controlled via compat.InitFromFlags. Grafana initializes the compat package without any feature flags, which is the equivalent of fallback mode. Classic and UTF-8 strict mode are used in Mimir. While Grafana Managed Alerts have no need for fallback mode, Grafana can still be used as an interface to manage the configurations of Mimir Alertmanagers and view configurations of Prometheus Alertmanager, and those installations might not have migrated or being running on older versions. Such installations behave as if in classic mode, and Grafana must be able to parse their configurations to interact with them for some period of time. As such, Grafana uses fallback mode until we are ready to drop support for outdated installations of Mimir and the Prometheus Alertmanager.	2024-02-06 08:33:47 +00:00
William Wernert	2ab7d3c725	Alerting: Receivers API (read only endpoints) (#81751 ) * Add single receiver method * Add receiver permissions * Add single/multi GET endpoints for receivers * Remove stable tag from time intervals See end of PR description here: https://github.com/grafana/grafana/pull/81672	2024-02-05 20:12:15 +02:00
Ryan McKinley	9c9e5e68c8	User: Add uid colum to user table (#81615 )	2024-02-01 18:14:10 -08:00
Yuri Tseretyan	d1073deefd	Alerting: Time intervals API (read only endpoints) (#81672 ) * declare new API and models GettableTimeIntervals, PostableTimeIntervals * add new actions alert.notifications.time-intervals:read and alert.notifications.time-intervals:write. * update existing alerting roles with the read action. Add to all alerting roles. * add integration tests	2024-02-01 15:17:13 -05:00
William Wernert	7e939401dc	Alerting: Introduce initial common receiver service (#81211 ) * Create locking config store that mimics existing provisioning store * Rename existing receivers(_test).go * Introduce shared receiver group service * Fix test * Move query model to models package * ReceiverGroup -> Receiver * Remove locking config store * Move convert methods to compat.go * Cleanup	2024-02-01 14:42:59 -05:00
George Robinson	0726c7c3fa	Alerting: Prevent inhibition rules in Grafana Alertmanager (#81712 ) This commit prevents saving configurations containing inhibition rules in Grafana Alertmanager. It does not reject inhibition rules when using external Alertmanagers, such as Mimir. This meant the validation had to be put in the MultiOrgAlertmanager instead of in the validation of PostableUserConfig. We can remove this when inhibition rules are supported in Grafana Managed Alerts.	2024-02-01 14:53:15 +00:00
Matthew Jacobson	0ce1ccd6f9	Alerting: Fix inconsistent AM raw config when applied via sync vs API (#81655 ) AM config applied via API would use the PostableUserConfig as the AM raw config and also the hash used to decide when the AM config has changed. However, when applied via the periodic sync the PostableApiAlertingConfig would be used instead. This leads to two issues: - Inconsistent hash comparisons when modifying the AM causing redundant applies. - GetStatus assumed the raw config was PostableUserConfig causing the endpoint to return correctly after a new config is applied via API and then nothing once the periodic sync runs. Note: Technically, the upstream GrafanaAlertamanger GetStatus shouldn't be returning PostableUserConfig or PostableApiAlertingConfig, but instead GettableStatus. However, this issue required changes elsewhere and is out of scope.	2024-01-31 21:05:30 +02:00
Ashley Harrison	39057552dc	QueryField: Handle autocomplete better (#81484 ) * extract out function + add unit tests * add feature toggle and default it to on	2024-01-31 10:01:20 +00:00
Yuri Tseretyan	131c72d655	Alerting: Fix scheduler to group folders by the unique key (orgID and UID) (#81303 )	2024-01-30 17:14:11 -05:00
Sofia Papagiannaki	89d3b55bec	Folders: Reduce DB queries when counting and deleting resources under folders (#81153 ) * Add folder store method for fetching all folder descendants * Modify GetDescendantCounts() to fetch folder descendants at once * Reduce DB calls when counting library panels under dashboard * Reduce DB calls when counting dashboards under folder * Reduce DB calls during folder delete * Modify folder registry to count/delete entities under multiple folders * Reduce DB calls when counting * Reduce DB calls when deleting	2024-01-30 18:26:34 +02:00
William Wernert	de662810cf	Alerting: Create instance of alert rule generator in historian annotation tests (#81394 ) * Create generator variable to ensure closures have correct context	2024-01-29 11:22:43 -05:00
idafurjes	f44592a97a	Remove folderID from service tests (#80615 ) * Remove folderID from service tests * Remove folderID from ngalert migration tests * Remove tests related to folderIDs * Roll back change Before removing FolderID from this test, we need to adjust the code * Remove FolderID from publicdashboard pkg * Add back annotations test	2024-01-26 17:36:35 +02:00
Gabriel MABILLE	722b78f3e0	RBAC: Add userLogin filter to the permission search endpoint (#81137 ) * RBAC: Search add user login filter * Switch to a userService resolving instead * Remove unused error * Fallback to use the cache * account for userID filter * Account for the error * snake case * Add test cases * Add api tests * Fix return on error * Re-order imports	2024-01-26 09:43:16 +01:00
Sofia Papagiannaki	b1eec36df3	Alerting: Fix authorisation to use namespace UIDs for scope (#81231 )	2024-01-25 15:19:51 -05:00
idafurjes	7e5544ab21	Add MFolderIDsServiceCount to count folderIDs in services pkg (#81237 )	2024-01-25 11:10:35 +01:00
Sofia Papagiannaki	478d7d58fa	Nested folders: Allow creating folders with duplicate names in different locations (#77076 ) * Add API test * Add move tests * Fix create folder * Fix move * Fix test * Drop and re-create index so that allows a folder to contain a dashboard and a subfolder with same name * Get folder by title defaults to root folder and optionally fetches folder by provided parent folder * Apply suggestions from code review	2024-01-25 11:29:56 +02:00
William Wernert	2203bc2a3d	Alerting: Refactor provisioning tests/fakes (#81205 ) * Fix up test Alertmanager config JSON * Move fake AM config and provisioning stores to fakes package	2024-01-24 17:15:55 -05:00
Matthew Jacobson	71e70c424f	Alerting: During legacy migration reduce the number of created silences (#78505 ) * Alerting: During legacy migration reduce the number of created silences During legacy migration every migrated rule was given a label rule_uid=<uid>. This was used to silence DatasourceError/DatasourceNoData alerts for migrated rules that had either ExecutionErrorState/NoDataState set to keep_state, respectively. This could potentially create a large amount of silences and a high cardinality label. Both of these scenarios have poor outcomes for CPU load and latency in unified alerting. Instead, this change creates one label per ExecutionErrorState/NoDataState when they are set to keep_state as well as two silence rules, if rules with said labels were created during migration. These silence rules are: - __legacy_silence_error_keep_state__ = true - __legacy_silence_nodata_keep_state__ = true This will drastically reduce the number of created silence rules in most cases as well as not create the potentially high cardinality label `rule_uid`.	2024-01-24 15:56:19 -05:00
Santiago	fbbda6c05e	Alerting: Retry readiness check to the remote Alertmanager on 5xx status code responses (#81174 )	2024-01-24 21:39:06 +01:00
George Robinson	05d858635c	Alerting: Add metric for inhibition rules (#81119 ) This commit adds a metric for the number of inhibition rules. It matches the metric added upstream in #3681.	2024-01-23 19:43:17 +00:00
Jean-Philippe Quéméner	aa25776f81	Alerting: Add a feature flag to periodically save states (#80987 )	2024-01-23 17:03:30 +01:00
George Robinson	85b9edcd28	Alerting: Fix incorrect initialization of logger (#81099 )	2024-01-23 17:29:38 +02:00
Marcus Efraimsson	6768c6c059	Chore: Remove public vars in setting package (#81018 ) Removes the public variable setting.SecretKey plus some other ones. Introduces some new functions for creating setting.Cfg.	2024-01-23 12:36:22 +01:00
Jean-Philippe Quéméner	eb7e1216a1	feat(alerting): add async state persister (#80763 )	2024-01-22 13:07:11 +01:00
Julien Duchesne	40312c527b	ngalert openapi: Fix ObjectMatchers definition (#79477 ) These don't get marshalled and unmarshalled in the same way as they are represented in Go This PR changes the OpenAPI spec to reflect what the API accepts and sends back	2024-01-19 14:37:11 -05:00
Alexander Weaver	18b9c8fd5f	Alerting: Nilcheck JitterStrategyFrom so it can be used in contexts without feature toggles (#80841 ) Nilcheck so tests can have a nil feature toggles	2024-01-18 15:43:41 -06:00
Alexander Weaver	00a260effa	Alerting: Add setting to distribute rule group evaluations over time (#80766 ) * Simple, per-base-interval jitter * Add log just for test purposes * Add strategy approach, allow choosing between group or rule * Add flag to jitter rules * Add second toggle for jittering within a group * Wire up toggles to strategy * Slightly improve comment ordering * Add tests for offset generation * Rename JitterStrategyFrom * Improve debug log message * Use grafana SDK labels rather than prometheus labels	2024-01-18 12:48:11 -06:00
Julien Duchesne	c9211fbd69	ngalert openapi: Use same `basePath` as rest of Grafana (#79025 ) * ngalert openapi: Use same `basePath` as rest of Grafana Currently, there are two issues that prevent easily merging `ngalert` and grafana openapi specs: - The basePath is different. `grafana` has `/api` and `ngalert` has `/api/v1`. I changed `ngalert` to use `/api` - The `ngalert` endpoints have their basePath in the each operation path. The basePath should actually be omitted --------- Co-authored-by: Yuriy Tseretyan <yuriy.tseretyan@grafana.com>	2024-01-17 11:53:16 -05:00
Jean-Philippe Quéméner	82638d059f	feat(alerting): add state persister interface (#80384 )	2024-01-17 13:33:13 +01:00
Santiago	3217a0dc05	Alerting: Fix state sync errors counter increment (#80702 )	2024-01-17 11:04:27 +01:00
Sofia Papagiannaki	d1dab5828d	Alerting: Update rule API to address folders by UID (#74600 ) * Change ruler API to expect the folder UID as namespace * Update example requests * Fix tests * Update swagger * Modify FIle field in /api/prometheus/grafana/api/v1/rules * Fix ruler export * Modify folder in responses to be formatted as <parent UID>/<title> * Add alerting test with nested folders * Apply suggestion from code review * Alerting: use folder UID instead of title in rule API (#77166) Co-authored-by: Sonia Aguilar <soniaaguilarpeiron@gmail.com> * Drop a few more latent uses of namespace_id * move getNamespaceKey to models package * switch GetAlertRulesForScheduling to use folder table * update GetAlertRulesForScheduling to return folder titles in format `parent_uid/title`. * fi tests * add tests for GetAlertRulesForScheduling when parent uid * fix integration tests after merge * fix test after merge * change format of the namespace to JSON array this is needed for forward compatibility, when we migrate to full paths * update EF code to decode nested folder --------- Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com> Co-authored-by: Virginia Cepeda <virginia.cepeda@grafana.com> Co-authored-by: Sonia Aguilar <soniaaguilarpeiron@gmail.com> Co-authored-by: Alex Weaver <weaver.alex.d@gmail.com> Co-authored-by: Gilles De Mey <gilles.de.mey@gmail.com>	2024-01-17 11:07:39 +02:00
Alexander Weaver	3c796ecc8f	Alerting: Add metric counting rule groups per org (#80669 ) * Refactor, fix bad map hint * Count groups per org	2024-01-16 16:35:56 -06:00
Santiago	3afd94185c	Alerting: Add metric to check for default AM configurations (#80225 ) * Alerting: Add metric to check for default AM configurations * Use a gauge for the config hash * don't go out of bounds when converting uint64 to float64 * expose metric for config hash * update metrics after applying config	2024-01-16 17:12:24 +01:00
Yuri Tseretyan	4b071f5452	Alerting: Fix MuteTiming Get API to return provenance status (#80494 )	2024-01-13 00:16:54 +02:00
Julien Duchesne	2fb03dfa56	fix(swagger): Mute Timing PUT OK status is 202 (#80459 )	2024-01-12 16:58:20 -05:00
Yuri Tseretyan	4479e7218d	Alerting: MuteTiming service return errutil + GetTiming by name (#79772 ) * add get mute timing by name to MuteTimingService * update get mute timing request handler to use the service method * replace validation, uniqueness and used errors with errutils * update mute timing methods return errutil responses * use the term "time interval" in errors bevause mute timings are deprecated in Alertmanager and will be replaced by time intervals in the future. * update create and update methods to return struct instead of pointer	2024-01-12 21:23:44 +02:00
idafurjes	cb419e799b	Remove folderid service test (#80433 ) * Remove FolderID from service tests * Add models * Add folderID pack to publicdashboard tests * Remove folderID from dashboard tests * Remove folderID from folders * Remove folderID from ngalert tests * Remove nolint comment * Add back some tests after rebase	2024-01-12 16:43:39 +01:00
Yuri Tseretyan	77db6a9ca4	Alerting: Fix GetAlertRulesForScheduling to use folder table and join by org_id (#80330 )	2024-01-11 09:21:03 -05:00
Santiago	6c87d9a1e7	Alerting: Stop retries on 4xx status code responses (remote Alertmanager readiness check) (#80350 )	2024-01-11 12:12:35 +01:00
William Wernert	48b5ac779b	Alerting/Annotations: Add annotation backend for Loki alert state history (#78156 ) * Move scope type vars to testutil package * Expose parts of state historian for use in annotation backend * Implement Loki ASH Annotation store This store will only implement the `Get` method of a RepositoryImpl since alert state history writes to Loki elsewhere. * Use interface for Loki HTTP Client * Add tests for Loki ASH Annotation store * Add missing test * Fix lint * Organize tests * Add filter tests * Improve tests * Move filter logic into outer function * Fix lint * Add comment * Fix tests * Fix lint * Rename historian store + refactor * Cleanup historian store * Fix tests * Minor cleanup * Use new `ShouldRecordAnnotation` filter * Fix logic and add tests for this check * Fix typos, remove unused variables, `< 1` -> `== 0` * More closely mimic RBAC filter from xorm to ensure correct logic * Move off weaveworks client * Address PR comments	2024-01-10 18:42:35 -05:00
Matthew Jacobson	afa33f12b2	Alerting: Create alertingQueryOptimization feature flag for alert query optimization (#78932 ) * Alerting: Create feature flag for alert query optimization Adds a feature flag alertingQueryOptimization for an already existing functionality: alert query optimization. This feature flag will now be disabled by default.	2024-01-10 15:52:58 -05:00
Matthew Jacobson	f365d35cf8	Alerting: Show warning when query optimized (#78751 ) * Alerting: Show warning when query optimized * Use frame.AppendNotices * Improve warning to include why and a prompt for action	2024-01-10 14:40:00 -05:00
Santiago	9e78faa7ba	Alerting: Add metrics to the remote Alertmanager struct (#79835 ) * Alerting: Add metrics to the remote Alertmanager struct * rephrase http_requests_failed description * make linter happy * remove unnecessary metrics * extract timed client to separate package * use histogram collector from dskit * remove weaveworks dependency * capture metrics for all requests to the remote Alertmanager (both clients) * use the timed client in the MimirAuthRoundTripper * HTTPRequestsDuration -> HTTPRequestDuration, clean up mimir client factory function * refactor * less git diff * gauge for last readiness check in seconds * initialize LastReadinesCheck to 0, tweak metric names and descriptions * add counters for sync attempts/errors * last config sync and last state sync timestamps (gauges) * change latency metric name * metric for remote Alertmanager mode * code review comments * move label constants to metrics package	2024-01-10 11:18:24 +01:00

1 2 3 4 5 ...

1393 Commits