grafana

mirror of https://github.com/grafana/grafana.git synced 2024-11-25 10:20:29 -06:00

Author	SHA1	Message	Date
William Wernert	c62cc25513	Alerting: Configure recording rule writer from config.ini (#89056 )	2024-06-12 16:04:46 -04:00
Santiago	b7120c5a30	Alerting: Fix missing argument in call to createRemoteAlertmanager() (#89101 )	2024-06-12 11:57:22 +03:00
Santiago	12d5251c12	Alerting: Alertmanager configuration sync loop (#88822 ) * make the config sync happen on each call to ApplyConfig(), fix tests * send autogen config * add fake autogen function for tests * update stale comments, tidy things up, make linter happy * add auto-gen routes only if the feature toggle is enabled * remove unnecessary fake autogen function * throttle configuration syncs * restore pkg/services/store/entity/sqlstash/sql_storage_server.go * test sync loop in ApplyConfig, skip invalid autogen routes * restore conf/defaults.ini * restore conf/defaults.ini * avoid skipping invalid auto-gen routes in SaveAndApplyConfig * test that autogenFn is called and its errors are returned * add debug message about the sync interval not having elapsed * collapse two log lines into one	2024-06-12 10:13:34 +02:00
Jacob Valdemar	eb76ea47a0	Alerting: Add ha_reconnect_timeout configuration option (#88823 ) * Docs: Update "Configure high availability" guide with ha_reconnect_timeout configuration --------- Co-authored-by: Christopher Moyer <35463610+chri2547@users.noreply.github.com>	2024-06-11 13:25:48 -04:00
Alexander Akhmetov	667fea6623	Alerting: use hash of labels instead of labels string as the alert state cache key (#88956 ) * Alerting: use hash instead of labels as the cache key * Use data.Labels.Fingerprint to calculate the cache key	2024-06-11 18:34:58 +02:00
Alexander Weaver	d004f8a98d	Alerting: Recording rules understands errors embedded in dataframes (#88946 ) * Make MakeDependencyError public for tests in another package * Create tests for errors in eval results * Extract logic to pull frame errors out into exported function * Maybe we can drop cyclomatic complexity lint suppression now? * extract frame errors and fail recording rules if frames contain error * Fix up retry logic to actually work * Do not retry non retryable errors	2024-06-11 10:37:10 -05:00
Steve Simpson	d440d86bbb	Alerting: Fix erroneous use of grafana-cli/logger. (#89037 ) Can't see how this was intentional, likely just a typo.	2024-06-11 14:29:56 +02:00
Santiago	cdbc9d801f	Alerting: Use the internal Alertmanager to test templates and receivers (remote primary) (#88988 )	2024-06-11 11:06:07 +02:00
Yuri Tseretyan	d4b0ac5973	Alerting: Fix rule storage to filter by group names using case-sensitive comparison (#88992 ) * add test for the bug * remove unused struct * update db store to post process filters by group using go-lang's case-sensitive string comparison -------- Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com>	2024-06-10 19:05:47 -04:00
Santiago	5f4d07bb75	Alerting: Enable remote primary mode using feature toggles (#88976 )	2024-06-10 17:07:13 +02:00
Santiago	e15e40fbd3	Alerting: Skip setting up clustering in remote primary/only modes (#88968 ) * Alerting: Skip setting up clustering in remote primary mode * Update pkg/services/ngalert/notifier/multiorg_alertmanager.go Co-authored-by: Steve Simpson <steve.simpson@grafana.com> --------- Co-authored-by: Steve Simpson <steve.simpson@grafana.com>	2024-06-10 13:51:11 +02:00
William Wernert	63e9969c1b	Alerting: Recording rule mapping logic for data frames to Prometheus metrics (#88550 ) * Add stub Prometheus writer with mapping logic * Add tests	2024-06-07 20:00:22 +03:00
Alexander Weaver	58fdb24b0b	Alerting: Recording rules appear as type=recording in Prometheus API + better abstraction for type (#88805 ) * Wire status through to prom API * Regenerate swagger	2024-06-07 11:24:06 -05:00
Yuri Tseretyan	32ea1801aa	Alerting: Support AWS SNS integration in Grafana (#88867 )	2024-06-07 11:49:49 -04:00
Alexander Weaver	f1dc63565e	Alerting: Fix go-swagger extraction and several embedded types from Alertmanager in Swagger docs (#88879 ) Drop redundant swagger model comments	2024-06-07 10:47:47 -05:00
Yuri Tseretyan	003e3efce9	Alerting: Update mute timings provisioning API to support optimistic locking (#88731 ) * add version to time-interval models * set time interval fingerprint as version * update to check provided version * delete to check if version is provided in query parameter 'version' * update integration tests * update specs	2024-06-06 18:06:37 -04:00
Alexander Weaver	a2e21d61f8	Alerting: Remove dead `evalRunning` guard in rule routine (#88312 ) Remove dead guard	2024-06-06 16:15:01 -05:00
William Wernert	d359591dac	Alerting: Support recording rule struct in provisioning API (#87849 ) * Support record struct in provisioning API * Update api spec * Use record field * Restrict API endpoints following toggle * Fix swagger spec * Add recording rule validation to store validator	2024-06-06 21:05:02 +03:00
Alexander Weaver	820ee6e9db	Alerting: Make all in api generator tooling now actually makes all (#88793 ) * Make all now actually makes all * Clean depends on clean-go	2024-06-05 11:52:31 -05:00
Fayzal Ghantiwala	80f54778f3	Alerting: Add option to use Redis in cluster mode for Alerting HA (#88696 ) * Add config option to use Redis in cluster mode * Use UniversalOptions	2024-06-05 17:02:25 +01:00
Dave Henderson	df784917e4	Alerting: Improve performance of tupleLablesToLabels function (#88736 ) * Alerting: Improve performance of tupleLablesToLabels function Signed-off-by: Dave Henderson <dave.henderson@grafana.com> * use %s for string rather than %v Co-authored-by: Sofia Papagiannaki <1632407+papagian@users.noreply.github.com> --------- Signed-off-by: Dave Henderson <dave.henderson@grafana.com> Co-authored-by: Sofia Papagiannaki <1632407+papagian@users.noreply.github.com>	2024-06-05 16:19:09 +03:00
Santiago	9f9928d41a	Alerting: Update grafana/alerting (#88363 ) * Alerting: Update grafana/alerting * make tests pass by implementing yaml unmarshallers and deleting fields with omitempty in their yaml tags * go mod tidy * fix tests by implementing not calling GettableApiAlertingConfig.UnmarshalYAML from GettableApiAlertingConfig.UnmarshalJSON * cleanup, reduce diff * fix more tests * update grafana/alerting to latest commit, delete global section from configs in tests * bring back YAML unmarshaller for GettableApiAlertingConfig * update alerting package dependency to point to main * skip test for sns notifier	2024-06-04 20:29:37 +02:00
Yuri Tseretyan	a63ef42816	Alerting: Mute Timing service to prevent changing provenance status to none (#88462 ) * use relaxed validation to not introduce breaking changes for now but to be able to use the service in non-provisioning APIs.	2024-06-04 08:54:33 -04:00
Fayzal Ghantiwala	b66cd7ef79	Alerting: Add filters for RouteGetRuleStatuses (#88295 ) * Placeholder commit with rule_uid change * Add new filters to grafana rule state API * Revert type change * Split rule_group and rule_name params * remove debug line * Change how query params are parsed * Comment	2024-06-04 10:57:55 +01:00
Matthew Jacobson	31d5dd0a12	Alerting: Prevent updating rule uid matcher for silences (#88519 ) Prevents updating the `__alert_rule_uid__` equality matcher (used for rule-specific silences) on existing silences	2024-06-03 17:39:06 -04:00
Fayzal Ghantiwala	67b9e3b269	Alerting: Update HA Redis TLS docs (#88538 ) * Update HA Redis TLS doc * Add test for regular TLS * Update docs * Update prom registry	2024-05-31 13:23:45 +01:00
Sofia Papagiannaki	17ca61d7f8	Alerting: Export and provisioning rules into subfolders (#77450 ) * Folders: Optionally include fullpath in service responses * Alerting: Export folder fullpath instead of title * Escape separator in folder title * Add support for provisiong alret rules into subfolders * Use FolderService for creating folders during provisioning * Export WithFullpath() folder service function --------- Co-authored-by: Tania B <yalyna.ts@gmail.com> Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-05-31 11:09:20 +03:00
Matthew Jacobson	09cb3a6048	Alerting: Add optional metadata via query param to silence GET requests (#88000 ) * Alerting: Add optional metadata to GET silence responses - ruleMetadata: to request rule metadata. - accesscontrol: to request access control metadata.	2024-05-30 12:04:47 -04:00
William Wernert	5de7d4d06d	Alerting: Create writer interface for recording rules (#88459 ) * Create writer interface for recording rules Also create fake impl + use it for stub in scheduler	2024-05-29 22:38:33 +03:00
Fayzal Ghantiwala	543f0ae37e	Alerting: Update ListAlertRulesQuery to take a slice of RuleGroups (#88385 ) * Change ListAlertRulesQuery to take RuleGroup slice instead * Change func name * Change func name * Fix fakes * Fix function arg	2024-05-29 11:50:33 +01:00
Alexander Weaver	b926b6336d	Alerting: Scheduled recording rules execute their queries (#88309 ) * Basic eval flow * Wiring-up * fix * Extend todo * Start with tests * Include some relevant tests, skip ones that seem to have timing-based race conditions * Some tests, touch up linter and todo * Solve TODO * Add tracing * Tests to make sure an eval went through * Wire up feature toggles * Update pkg/services/ngalert/schedule/recording_rule.go Co-authored-by: Steve Simpson <steve.simpson@grafana.com> * Update pkg/services/ngalert/schedule/recording_rule_test.go Co-authored-by: Steve Simpson <steve.simpson@grafana.com> * Update pkg/services/ngalert/schedule/recording_rule_test.go Co-authored-by: Steve Simpson <steve.simpson@grafana.com> * Update pkg/services/ngalert/schedule/recording_rule_test.go Co-authored-by: Steve Simpson <steve.simpson@grafana.com> --------- Co-authored-by: Steve Simpson <steve.simpson@grafana.com>	2024-05-28 10:59:21 -05:00
Matthew Jacobson	8418aca823	Alerting: Add single rule checks to alert rule access control (#88307 ) * Alerting: Add single rule checks to alert rule access control Modifies ruler api single rule read to no longer fetch entire groups and instead use the new single rule ac check. Simplifies provisioning api getAlertRuleAuthorized logic to always load a single rule instead of conditionally loading the entire group when provisioning permissions are not present. * Swap out Has/AuthorizeAccessToRule for Has/AuthorizeAccessInFolder	2024-05-28 10:49:24 -04:00
Kyle Brandt	a738cb42d8	Prometheus: Update dependency to v0.52.0 (#87809 ) * Prometheus: Update dependency to v0.52.0 * go work sync * fix panics in tests * go work sync * prometheus v0.52.0 * handle errors * Update pkg/services/ngalert/sender/sender_test.go Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> * Update pkg/services/ngalert/sender/sender_test.go Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> --------- Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> Co-authored-by: ismail simsek <ismailsimsek09@gmail.com>	2024-05-28 15:22:20 +02:00
Steve Simpson	08b18113d2	Alerting: Wire up alertmanagerRemoteOnly feature toggle. (#88329 ) * Alerting: Wire up alertmanagerRemoteOnly feature toggle. Though the mode isn't feature complete yet, it will be useful to have the feature toggle wired up in order to start testing. * Apply suggestions from code review Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> * Formatting --------- Co-authored-by: Santiago <santiagohernandez.1997@gmail.com>	2024-05-27 16:18:46 +02:00
Steve Simpson	ed42119907	Alerting: Pass metrics Registerer into NewExternalAlertmanagerSender. (#88313 ) * Alerting: Pass metrics Registerer into NewExternalAlertmanagerSender. I will work on a separate change to export the metrics from Grafana, this is a little more complicated. * Typo	2024-05-24 23:03:34 +02:00
Steve Simpson	8bcf589301	Alerting: Pass logger into NewExternalAlertmanagerSender (#88256 )	2024-05-24 20:11:26 +02:00
Alexander Weaver	65793440d3	Alerting: Test infrastructure for recording rules (#88200 ) * Add test rule generator support for recording rules * Remove accidental add * Recording rules appear in GetRulesForScheduling * A couple more tests, updates, count * No need to capture rule defs	2024-05-23 16:27:07 -05:00
William Wernert	006d0021e3	Alerting: Remove requirement for datasource query on rule read (#87349 ) * Remove requirement for datasource query for rule read * Address PR comments	2024-05-23 12:44:30 -04:00
Matthew Jacobson	bc5d077b30	Alerting: separate out silence auth service preconditions checks (#87998 ) * Alerting: separate out silence auth service preconditions checks Will be useful for subsequent PR that adds metadata to silence response * Add silence read wildcard scope to precondition for read all silences	2024-05-23 12:34:42 -04:00
Steve Simpson	8421919cb5	Alerting: Feature toggle to disallow sending alerts externally (#87982 ) * Define feature toggle * Implement feature toggle	2024-05-23 14:29:19 +02:00
Gaurav Agrawal	fdaa091a4d	Alerting: Support custom API URL for PagerDuty integration (#88007 ) * fix assert in LINE * fix pagerduty asserts --------- Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-05-22 15:31:55 -04:00
Alexander Weaver	89b54d06e9	Alerting: Schedule a shim implementation for recording rules (#87939 ) * Add shim rule implementation for recording rules * Give ruleFactory access to the original rule definition * Schedule shim implementation if the rule is a recording rule * Fix or suppress linter * Fix nolint	2024-05-21 16:42:58 -05:00
Alexander Weaver	49c8deb1ea	Alerting: Add recording rules to ruler API and validation (#87779 ) * Read path, main API * Define record field for incoming requests * Refactor several alerting specific validators into two paths * Refactor validateCondition actually contain all the condition validation logic * Move condition validation inside rule path * Validators for recording rules * Wire feature flag through to validators * Test for accepting a valid recording rule * Tests for negative case, no UID * Test for ignoring alerting fields * Build conditions based on recording rules as well * Regenerate swagger docs * Fix CRUD test to cover the right thing * Re-generate swagger docs with backdated v0.30.2 version * Regenerate base spec * Regenerate ngalert specs * Regenerate top level specs * Comment and rename * Return struct instead of modifying ref	2024-05-21 14:39:28 -05:00
William Wernert	cb0bcb6fe4	Alerting: Fix/update alerting API spec (#88130 )	2024-05-21 10:06:44 -04:00
Santiago	60e7a4e746	Alerting/Chore: Remove unused parameters (#88045 ) Alerting/Chore: Remove unused parameters from redisPeer.receiveLoop() and ReceiverService.shouldDecrypt()	2024-05-20 16:37:39 +02:00
Yuri Tseretyan	8c2a382788	Alerting: Fix typo in JSON response for rule export. (#88028 )	2024-05-20 09:39:39 -04:00
Yuri Tseretyan	05d6813a09	Alerting: Fix scheduler to sort rules before evaluation (#88006 ) sort rules scheduled for evaluation to make sure that the order is stable between evaluations. This is especially important in HA mode.	2024-05-17 11:38:19 -04:00
Santiago	e41434c332	Alerting: Promote configuration in the remote Alertmanager (#87388 )	2024-05-16 12:06:03 +02:00
Yuri Tseretyan	f410c7fca1	Alerting: use logger with same context within rule scheduling loop (#87934 )	2024-05-15 15:38:00 -04:00
Alexander Weaver	1badcf4b63	Alerting: Allow NoData and ExecErrState to be fully blank on recording rules (#87868 ) * Allow empty NoData and ExecErrState on recording rules * remove TODO about this	2024-05-15 09:35:54 -05:00
Alexander Weaver	b8a284fb81	Alerting: Fix xorm serialization of Record field struct, add tests for storing and reading (#87857 ) Fix sub struct ser and deser, add tests	2024-05-14 14:50:06 -05:00
Steve Simpson	67fa96f88d	Alerting: Pass logger into NewAnnotationBackend. (#87812 ) * Alerting: Pass logger into NewAnnotationBackend. Make it possible to pass loggers into more places for code reuse. * Mistake in passing logger	2024-05-14 15:51:27 +02:00
William Wernert	563fcb8bf4	Alerting: Encode query model map to string in rule export to avoid html escape sequences (#87663 ) * Encode query model map to string to avoid html escape sequences * Remove insignificant whitespace in test request	2024-05-14 09:29:50 -04:00
Fayzal Ghantiwala	7a2fbad0c8	Alerting: Add options to configure TLS for HA using Redis (#87567 ) * Add Alerting HA Redis Client TLS configs * Add test to ping miniredis with mTLS * Update .ini files and docs * Add tests for unified alerting ha redis TLS settings * Fix malformed go.sum * Add modowner * Fix lint error * Update docs and use dstls config	2024-05-14 14:21:42 +01:00
Alexander Weaver	e39658097f	Alerting: Wire recording rules feature toggle into limits struct (#87778 ) Wire toggle into limits	2024-05-14 07:44:14 -05:00
Ieva	167151b211	Chore: Remove use of deprecated method in AC code (#87541 ) * switch from using cfg to using featuremgmt for checking a feature toggle in AC code * merge test fixes	2024-05-10 11:56:52 +01:00
Alexander Weaver	a6a9ab4008	Alerting: Do not store series values from past evaluations in state manager for no reason (#87525 ) Do not store previous execution results on states	2024-05-09 15:51:55 -05:00
Yuri Tseretyan	356a29592b	Alerting: Add two sets of provisioning actions for rules and notifications (#87149 )	2024-05-09 13:19:07 -04:00
Alexander Weaver	36ef611cf4	Alerting: Add database migration for recording rule fields (#87012 ) * Create recording rule fields in model * Add migration * Write to database, support in version table * extend fingerprint * Force fields to be empty on validate * Another storage spot, tests for fingerprint * Explicitly set defaults in provisioning API * Tests for main API validation * Add diff tests even though fields are unpopulated for now * Use struct tag approach instead of FromDB/ToDB hooks as it better handles nulls when deserializing * test for deser * Backout RecordTo for now since it's not decided in the doc * back out of migration too * Drop datasourceref for now * address linter complaints * Try a single outer struct with all fields embedded	2024-05-09 12:12:44 -05:00
Alexander Weaver	6c47968f6c	Alerting: Do not retry rule evaluations with "input data must be a wide series but got type long" style errors (#87343 ) add typed error for series must be wide, do not retry	2024-05-07 11:31:07 -05:00
Matthew Jacobson	babfa2beac	Alerting: Hook up GMA silence APIs to new authentication handler (#86625 ) This PR connects the new RBAC authentication service to existing alertmanager API silence endpoints.	2024-05-03 15:32:30 -04:00
Santiago	b76a9e4d31	Alerting: Implement GetStatus in the remote Alertmanager struct (#84887 ) * Alerting: Implement GetStatus in the remote Alertmanager struct * update tests * fix tests, extract AlertmanagerConfig from PostableConfig * get the remote AM config instead of the Grafana one from the remote AM * pass grafana AM config in test * return error in GetStatus instead of logging it (internal AM)	2024-05-03 13:59:02 +02:00
Fayzal Ghantiwala	df25e9197e	Alerting: Get grafana-managed alert rule by UID (#86845 ) * Add auth checks and test * Check user is authorized to view rule and add tests * Change naming * Update Swagger params * Update auth test and swagger gen * Update swagger gen * Change response to GettableExtendedRuleNode * openapi3-gen * Update tests with refactors models pkg	2024-05-02 15:24:59 +01:00
Serge Zaitsev	ad5613d7d4	Chore: Remove cfg from folder service (#87212 ) remove cfg from folder service	2024-05-02 13:18:54 +02:00
William Wernert	93519f70ca	Alerting: Also fix HCL field name for MuteTimeIntervals (#87079 ) * Correct HCL field name for MuteTimeIntervals * Update test	2024-04-30 16:14:01 +01:00
Yuri Tseretyan	052082a927	Alerting: Refactor Alert Rule Generators (#86813 )	2024-04-29 21:52:15 -04:00
William Wernert	70ff229bed	Alerting: Use expected field name for receiver in HCL export (#87065 ) * Use expected field name for receiver in hcl Terraform provider expects `contact_point` instead of `receiver` in notification settings on a rule.	2024-04-29 18:13:29 +01:00
Santiago	36a0499128	Alerting: Implement CreateSilence in the forked Alertmanager (remote primary mode) (#85716 )	2024-04-29 18:47:25 +02:00
Santiago	1af2e69625	Alerting: Implement DeleteSilence in the forked AM (remote primary) (#85721 )	2024-04-29 17:23:41 +02:00
Steve Simpson	fbaa847a3c	Alerting: Pass logger into NewRemoteLokiBackend. (#87029 ) Tiny refactor to allow a logger to be passed into NewRemoteLokiBackend.	2024-04-29 12:10:23 +02:00
Yuri Tseretyan	dff7cb9afb	Alerting: Move alertmanager api silence code to separate files (#86947 ) * Move alertmanager api silence code to separate files unchanged * Replace with silence model instead interface --------- Co-authored-by: Matt Jacobson <matthew.jacobson@grafana.com>	2024-04-25 15:20:37 -04:00
Matthew Jacobson	3397e8bf09	Alerting: Improve error when receiver or time interval used by rule is deleted (#86865 ) * Alerting: Improve error when receiver used by rule is deleted * Remove RuleUID from public error and data * Improve fallback error in am config post * Refactor to expand to time intervals * Fix message on unchecked errors to be same as before	2024-04-25 13:36:00 -04:00
Santiago	a6be12c037	Alerting: Implement SaveAndApplyConfig in the forked Alertmanager (remote primary) (#84659 ) * Alerting: Implement SaveAndApplyConfiguration in the forked Alertmanager struct * call SaveAndApplyConfig on the remote first, log errors for the internal * add comments explaining why we ignore errors in the internal AM * restore go.work.sum	2024-04-23 15:45:35 +02:00
Steve Simpson	a6ad2380bf	Alerting: Refactor api_prometheus.go request handlers. (#86639 ) This splits the request handlers into two functions, one which is the actual handler and one which is independent from the Grafana `ReqContext` object. This is to make it easier to reuse the implementation in other code. Part of the refactoring changes the functions which get query parameters from the request to operate on a `url.Values` instead of the request object. The change also makes the code consistently use `req.Form` instead of a combination of `req.URL.Query()` and `req.Form`, though I have left `api_ruler` as-is to avoid this PR growing too large.	2024-04-23 14:50:26 +02:00
Santiago	c77ab53819	Alerting: implement SaveAndApplyConfig in the remote Alertmanager struct (#84642 ) * implement SaveAndApplyConfig in the remote Alertmanager struct * remove ID from CreateGrafanaAlertmanagerConfig call * decrypt, test that we decrypt, refactor * fix duplicated declaration in test * rephrase comment, remove unnecessary conversion to slice of bytes * fix test	2024-04-23 14:37:10 +02:00
Santiago	8b7c2a459b	Alerting: Implement SaveAndApplyDefaultConfig in the forked Alertmanager (remote primary mode) (#85668 ) * Alerting: Implement SaveAndApplyDefaultConfig in the forked Alertmanager (remote primary) * log the error for the internal AM instead of returning it	2024-04-23 14:36:40 +02:00
Yuri Tseretyan	9735a8a080	Alerting: Distinguish conflict violation errors (#86634 ) * update generator to set ID = 0 and do not set 0 if unique is needed * return proper message when the constraint violation	2024-04-22 12:28:46 -04:00
Julian Siebert	14f018e3fc	Docs: Use correct description for "og_priority" (#80889 )	2024-04-22 13:53:18 +00:00
Steve Simpson	54290f2ac4	Alerting: Fix TestRouteGetRuleStatuses as much as possible. (#86666 ) This test has been skipped for a long time, so it doesn't work anymore. I've fixed the test so it works again, but left some tests disabled which were apparently flaky. If we see the other test cases flaking, we'll have to disable it again. Fixes: - Use fake access control for most test cases, and real one for FGAC test cases. - Check that "file" in API responses the full folder path, not folder title.	2024-04-22 12:36:50 +02:00
Steve Simpson	f07f48616a	Alerting: Fix panic when limit_alerts=0. (#86640 ) Oversight in the TopK function meant if k=0, then we'd panic when checking element zero in the heap, because no items are ever allowed into the heap.	2024-04-22 10:14:19 +02:00
Steve Simpson	6ea97e41fb	Alerting: Consistently return Prometheus-style responses from rules APIs. (#86600 ) * Alerting: Consistently return Prometheus-style responses from rules APIs. This commit is part refactor and part fix. The /rules API occasionally returns error responses which are inconsistent with other error responses. This fixes that, and adds a function to map from Prometheus error type and HTTP code. * Fix integration tests * Linter happiness * Make linter more happy * Fix up one more place returning non-Prometheus responses	2024-04-19 21:03:20 +02:00
Santiago	529f55cfe8	Alerting: Remove isDefault field from receivers (Alertmanager configuration) (#86605 ) Alerting: Remove isDefault field from receivers in the Alertmanager configuration	2024-04-19 15:44:20 +02:00
Santiago	309a7e7684	Alerting: Implement SaveAndApplyDefaultConfig in the remote Alertmanager struct (#85005 ) * Alerting: Implement SaveAndApplyDefaultConfig in the remote Alertmanager struct * send the hash of the encrypted configuration * tests, default config hash in AM struct * add missing default config to test * restore build directory * go work file... * fix broken test * remove unnecessary conversion to []byte * go work again... * make things work again with latest main branch changes * update error messages in tests for decrypting config	2024-04-19 15:11:07 +02:00
Santiago	a2ce8fefed	Alerting: Use a struct when sending a Grafana AM configuration to the remote Alertmanager (#86451 ) * Alerting: Use a struct when sending a Grafana AM configuration to the remote Alertmanager * remove '-distroless' from mimir image name	2024-04-19 13:04:18 +02:00
Steve Simpson	5f7612834e	Alerting: Refactoring in api_prometheus.go to allow code reuse. (#86575 ) Preparing these functions to be used by some other part of the codebase, which does not have a `contextmodel.ReqContext`, only the normal request structure (`url.Values`, etc). This is slightly messy because of how Grafana allows url parameters to be in the URL or in the request body, so we need to make sure to invoke the form parsing logic in `ReqContext`.	2024-04-19 12:52:01 +02:00
Steve Simpson	73873f5a8a	Alerting: Optimize rule status gathering APIs when a limit is applied. (#86568 ) * Alerting: Optimize rule status gathering APIs when a limit is applied. The frontend very commonly calls the `/rules` API with `limit_alerts=16`. When there are a very large number of alert instances present, this API is quite slow to respond, and profiling suggests that a big part of the problem is sorting the alerts by importance, in order to select the first 16. This changes the application of the limit to use a more efficient heap-based top-k algorithm. This maintains a slice of only the highest ranked items whilst iterating the full set of alert instances, which substantially reduces the number of comparisons needed. This is particularly effective, as the `AlertsByImportance` comparison is quite complex. I've included a benchmark to compare the new TopK function to the existing Sort/limit strategy. It shows that for small limits, the new approach is much faster, especially at high numbers of alerts, e.g. 100K alerts / limit 16: 1.91s vs 0.02s (-99%) For situations where there is no effective limit, sorting is marginally faster, therefore in the API implementation, if there is either a) no limit or b) no effective limit, then we just sort the alerts as before. There is also a space overhead using a heap which would matter for large limits. * Remove commented test cases * Make linter happy	2024-04-19 11:51:22 +02:00
Matthew Jacobson	a20197229e	Alerting: Prevent simplified routing zero duration GroupInterval and RepeatInterval (#86561 ) Prevent zero duration GroupInterval and RepeatInterval	2024-04-18 21:08:38 -04:00
Matthew Jacobson	71445002b7	Alerting: Fix simplified routing group by override (#86552 ) * Alerting: Fix simplified routing custom group by override Custom group by overrides for simplified routing were missing required fields GroupBy and GroupByAll normally set during upstream Route validation. This fix ensures those missing fields are applied to the generated routes. * Inline GroupBy and GroupByAll initialization instead of normalize after	2024-04-18 21:08:14 -04:00
Matthew Jacobson	533bed6d94	Alerting: Fix simplified routes '...' groupBy creating invalid routes (#86006 ) * Alerting: Fix simplified routes '...' groupBy creating invalid routes There were a few ways to go about this fix: 1. Modifying our copy of upstream validation to allow this 2. Modify our notification settings validation to prevent this 3. Normalize group by on save 4. Normalized group by on generate Option 4. was chosen as the others have a mix of the following cons: - Generated routes risk being incompatible with upstream/remote AM - Awkward FE UX when using '...' - Rule definition changing after save and potential pitfalls with TF With option 4. generated routes stay compatible with external/remote AMs, FE doesn't need to change as we allow mixed '...' and custom label groupBys, and settings we save to db are the same ones requested. In addition, it has the slight benefit of allowing us to hide the internal implementation details of `alertname, grafana_folder` from the user in the future, since we don't need to send them with every FE or TF request. * Safer use of DefaultNotificationSettingsGroupBy * Fix missed API tests	2024-04-16 12:14:39 -04:00
Alexander Weaver	5b1498f98f	Alerting: Return a 400 and errutil error when trying to delete a contact point that is referenced by a policy (#85481 ) Return a 400 and errutil error when trying to delete a contact point that is referenced by a policy	2024-04-15 09:25:28 -05:00
Yuri Tseretyan	12605bfed2	Alerting: Update fixed roles to include silences permissions (#85826 ) * update fixed roles to include silences * add silence actions to managed permissions * update documentation	2024-04-12 12:37:34 -04:00
Steve Simpson	ad7f804255	Alerting: Fix evaluation metrics to not count retries (#85873 ) * Change evaluation metrics to only count once per eval, and add new metrics. * Cosmetic: Move eval total Inc() to orginal place.	2024-04-12 16:20:46 +02:00
Matthew Jacobson	f79dd7c7f9	Alerting: Persist silence state immediately on Create/Delete (#84705 ) * Alerting: Persist silence state immediately on Create/Delete Persists the silence state to the kvstore immediately instead of waiting for the next maintenance run. This is used after Create/Delete to prevent silences from being lost when a new Alertmanager is started before the state has persisted. This can happen, for example, in a rolling deployment scenario. * Fix test that requires real data * Don't error if silence state persist fails, maintenance will correct	2024-04-09 13:39:34 -04:00
Santiago	2e7cc68394	Alerting: Remove CleanUp method from the Alertmanager (#85650 ) Alerting: Remove Cleanup method from the Alertmanager	2024-04-09 12:13:27 +02:00
Yuri Tseretyan	509691b416	Alerting: Introduce authorization logic for operations on silences (#85418 ) * extract genericService from RuleService just to reuse it later * implement silence service --------- Co-authored-by: William Wernert <william.wernert@grafana.com> Co-authored-by: Matthew Jacobson <matthew.jacobson@grafana.com>	2024-04-08 18:02:28 -04:00
Santiago	6a75a8f354	Alerting: Update grafana/alerting and use Upsert for creating silences (#85676 ) * Alerting: Update grafana/alerting and use Upsert for creating silences * go.work.sum * change error message in tests for silences (save -> upsert)	2024-04-08 11:46:14 +02:00
Alexander Weaver	03114e7602	Alerting: Return better error for invalid time range on alert queries (#85611 ) * Return better error for invalid time range * drop comment	2024-04-05 09:20:21 -05:00
Alexander Weaver	734d0111cb	Alerting: Export pure function to convert query results to alert results (#85393 ) Exported pure function to convert query results to alert results	2024-04-05 08:57:31 -05:00
Santiago	c7573bb0f7	Alerting: Make retention period configurable for the notification log (#85605 ) * Alerting: Make retention period configurable for the notification log * update sample.ini * fix outdated comment (on disk -> kvstore) * skip checking cyclomatic complexity for ReadUnifiedAlertingSettings	2024-04-05 12:25:43 +02:00
Alexander Weaver	623ee3a2be	Alerting: Only append `/alertmanager` when sending alerts to mimir targets if not already present (#85543 ) Don't append alertmanager if not present	2024-04-04 11:58:41 -05:00
Dave Henderson	5687243d0b	Feature Flags: use FeatureToggles interface where possible (#85131 ) * Feature Flags: use FeatureToggles interface where possible Signed-off-by: Dave Henderson <dave.henderson@grafana.com> * Replace TestFeatureToggles with existing WithFeatures Signed-off-by: Dave Henderson <dave.henderson@grafana.com> --------- Signed-off-by: Dave Henderson <dave.henderson@grafana.com>	2024-04-04 12:22:31 -04:00
Serge Zaitsev	faa1244518	Chore: Replace sqlstore with db interface (#85366 ) * replace sqlstore with db interface in a few packages * remove from stats * remove sqlstore in admin test * remove sqlstore from api plugin tests * fix another createUser * remove sqlstore in publicdashboards * remove sqlstore from orgs * clean up orguser test * more clean up in sso * clean up service accounts * further cleanup * more cleanup in accesscontrol * last cleanup in accesscontrol * clean up teams * more removals * split cfg from db in testenv * few remaining fixes * fix test with bus * pass cfg for testing inside db as an option * set query retries when no opts provided * revert golden test data * rebase and rollback	2024-04-04 15:04:47 +02:00
Jean-Philippe Quéméner	7cfd470c91	fix(alerting): only expose metrics if executing alerts (#85512 )	2024-04-03 17:18:02 +02:00
Benoit Tigeot	6f38ac6615	Alerting: Reduce set of fields that could trigger alert state change (#83496 ) We want to avoid too much change of alert state based on change on alert's fields. For that we ignore some fields from the diff.	2024-03-26 12:35:30 -04:00
Julien Duchesne	2188516a21	Alerting: Fix receiver inheritance when provisioning a notification policy (#82007 ) Terraform Issue: grafana/terraform-provider-grafana#1007 Nested routes should be allowed to inherit the contact point from the root (or direct parent) route but this fails in the provisioning API (it works in the UI)	2024-03-26 12:31:59 -04:00
ismail simsek	6137c4e0a6	Chore: Bump golangci-lint v1.57.1 (#84998 ) * bump golangci-lint v1.57.1 * update setting * remove goconst * fix linting issues * prettier * fix G601 * go mod tidy go work sync	2024-03-25 15:28:24 +01:00
Matthew Jacobson	0c3c5c5607	Alerting: Stop persisting silences and nflog to disk (#84706 ) With this change, we no longer need to persist silence/nflog states to disk in addition to the kvstore	2024-03-23 00:37:33 +02:00
Yuri Tseretyan	48de8657c9	Alerting: Editor role can access all provisioning API (#85022 )	2024-03-23 00:14:15 +02:00
Yuri Tseretyan	b9abb8cabb	Alerting: Update provisioning API to support regular permissions (#77007 ) * allow users with regular actions access provisioning API paths * update methods that read rules skip new authorization logic if user CanReadAllRules to avoid performance impact on file-provisioning update all methods to accept identity.Requester that contains all permissions and is required by access control. * create deltas for single rul e * update modify methods skip new authorization logic if user CanWriteAllRules to avoid performance impact on file-provisioning update all methods to accept identity.Requester that contains all permissions and is required by access control. * implement RuleAccessControlService in provisioning * update file provisioning user to have all permissions to bypass authz * update provisioning API to return errutil errors correctly --------- Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com>	2024-03-22 15:37:10 -04:00
Yuri Tseretyan	e138ae3eb9	Alerting: Improve openAPI specification and docs for export endpoints (#85008 )	2024-03-22 18:25:27 +02:00
Jean-Philippe Quéméner	f2c7023fe6	fix(alerting): use uid and not rand() in tests for title (#85001 )	2024-03-22 16:26:09 +02:00
Santiago	a2facbecd4	Alerting: Implement ApplyConfig for remote primary mode (forked AM) (#84811 ) * Alerting: Implement ApplyConfig for remote primary mode (forked AM) * add TODO for saving the config hash in other config-related methods * fix bad method receiver name (m -> am) * tests * add mutex * remove sync loop	2024-03-22 15:17:41 +01:00
Pepe Cano	2d6586952d	Alerting: Add placeholder to the Email Contact Point Message (#84064 )	2024-03-21 13:03:12 -04:00
Matthew Jacobson	fbd057b258	Alerting: Stop returning autogen routes for non-admin on api/v2/status (#84864 ) * Alerting: Stop returning autogen routes for non-admin on api/v2/status * Improve api/v2/status integration tests for user roles	2024-03-20 22:04:35 +02:00
William Wernert	6d16cf2699	Alerting: Marshal incoming json.RawMessage in diff (#84692 ) This will ensure the encoding is correct when comparing to the existing rule.	2024-03-20 13:10:39 -04:00
Yuri Tseretyan	04c9f459ec	Alerting: do not check for folder in file provisioning (#84822 ) provide nil folder service in file provisioning	2024-03-20 10:39:03 -04:00
Yuri Tseretyan	e593d36ed8	Alerting: Update rule access control to explicitly check for permissions "alert.rules:read" and "folders:read" (#78289 ) * require "folders:read" and "alert.rules:read" in all rules API requests (write and read). * add check for permissions "folders:read" and "alert.rules:read" to AuthorizeAccessToRuleGroup and HasAccessToRuleGroup * check only access to datasource in rule testing API --------- Co-authored-by: William Wernert <william.wernert@grafana.com>	2024-03-19 22:20:30 -04:00
Yuri Tseretyan	9dc4221508	Alerting: Log expression command types during evaluation (#84614 )	2024-03-19 10:00:03 -04:00
Santiago	4ad6d66479	Alerting: Remove ID from UserGrafanaConfig struct (#84602 ) * Alerting: Remove ID from UserGrafanaConfig struct * user custom mimir image withoud id in grafana config * change mimir image name	2024-03-19 12:47:13 +01:00
Santiago	c9bb18101c	Alerting: Decrypt secrets before sending configuration to the remote Alertmanager (#83640 ) * (WIP) Alerting: Decrypt secrets before sending configuration to the remote Alertmanager * refactor, fix tests * test decrypting secrets * tidy up * test SendConfiguration, quote keys, refactor tests * make linter happy * decrypt configuration before comparing * copy configuration struct before decrypting * reduce diff in TestCompareAndSendConfiguration * clean up remote/alertmanager.go * make linter happy * avoid serializing into JSON to copy struct * codeowners	2024-03-19 12:12:03 +01:00
Matthew Jacobson	3ea5c08c88	Alerting: External AM fix parsing basic auth with escape characters (#84681 )	2024-03-18 13:04:57 -04:00
William Wernert	97f37b2e6f	Alerting: Clamp Loki ASH range query to configured max_query_length (#83986 ) * Clamp range in loki http client to configured max_query_length Defaults to 721h to match Loki default	2024-03-15 18:59:45 +02:00
Yuri Tseretyan	827860d459	Alerting: Alerting accesscontrol utilities (#84508 ) * create fake for accesscontrol.RuleService * make errAuthorizationGeneric public	2024-03-14 14:03:53 -04:00
Yuri Tseretyan	f7d836feed	Alerting: Update rule provisioning service to accept user (#84480 )	2024-03-14 12:04:10 -04:00
Gilles De Mey	8765c48389	Alerting: Remove legacy alerting (#83671 ) Removes legacy alerting, so long and thanks for all the fish! 🐟 --------- Co-authored-by: Matthew Jacobson <matthew.jacobson@grafana.com> Co-authored-by: Sonia Aguilar <soniaAguilarPeiron@users.noreply.github.com> Co-authored-by: Armand Grillet <armandgrillet@users.noreply.github.com> Co-authored-by: William Wernert <rwwiv@users.noreply.github.com> Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-03-14 15:36:35 +01:00
William Wernert	8690a42e33	Alerting: Disallow invalid rule namespace UIDs in provisioning API (#83938 ) * Disallow invalid rule namespace UIDs in provisioning Reject requests with rules that reference a nonexistent folder or have an empty folder uid	2024-03-14 09:58:25 -04:00
Yuri Tseretyan	cfc3957894	Alerting: move store.ErrAlertRuleGroupNotFound to models package (#84308 ) move ErrAlertRuleGroupNotFound to models to avoid future circular dependencies	2024-03-12 15:38:21 -04:00
Andres Martinez Gotor	265200799d	Chore: Update grafana-plugin-sdk (#84289 )	2024-03-12 17:13:23 +01:00
William Wernert	10dc6c6d75	Alerting: Add "Keep Last State" backend functionality (#83940 ) * Implement keep last state for state transitions * Respect For duration when keeping state * Only keep transition from recording an annotation * Add keep last state option for nodata/error in UI	2024-03-12 10:00:43 -04:00
Alexander Weaver	6c5e94095d	Alerting: Scheduler and registry handle rules by an interface (#84044 ) * export Evaluation * Export Evaluation * Export RuleVersionAndPauseStatus * export Eval, create interface * Export update and add to interface * Export Stop and Run and add to interface * Registry and scheduler use rule by interface and not concrete type * Update factory to use interface, update tests to work over public API rather than writing to channels directly * Rename map in registry * Rename getOrCreateInfo to not reference a specific implementation * Genericize alertRuleInfoRegistry into ruleRegistry * Rename alertRuleInfo to alertRule * Comments on interface * Update pkg/services/ngalert/schedule/schedule.go Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com> --------- Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com>	2024-03-11 22:57:38 +02:00
Alexander Weaver	201f5d3ac9	Alerting: Extract large closures in ruleRoutine (#84035 ) * extract notify * extract resetState * move evaluate metrics inside evaluate * split out evaluate	2024-03-06 16:39:23 -06:00
Alexander Weaver	7a171fd14a	Regenerate openapidocs at 1.21.8 to match ci (#84037 ) * Regenerate openapidocs at 1.21.8 to match ci * Adjust trigger to work on the actual outputted files * Also put go.mod and go.sum in the triggers * manually fix * Make an arbitrary change rather than touching the trigger to force a run * Drop all triggers - run all the time * Print diff - taken from @papagian's PR * Manual fixes to swagger doc --------- Co-authored-by: Ryan McKinley <ryantxu@gmail.com>	2024-03-06 16:08:45 -06:00
gotjosh	948c8c45d6	Alerting: Use Alertmanager types extracted into grafana/alerting (#83824 ) * Alerting: Use Alertmanager types extracted into grafana/alerting We're in the process of exporting all Alertmanager types into grafana/alerting so that they can be imported in the Mimir Alertmanager, without a neeed to import Grafana directly. This change introduces type aliasing for all Alertmanager types based on their 1:1 copy that now live in grafana/alerting. Signed-off-by: gotjosh <josue.abreu@gmail.com> --------- Signed-off-by: gotjosh <josue.abreu@gmail.com>	2024-03-06 20:48:32 +00:00
Alexander Weaver	d5fda06147	Alerting: Decouple rule routine from scheduler (#84018 ) * create rule factory for more complicated dep injection into rules * Rules get direct access to metrics, logs, traces utilities, use factory in tests * Use clock internal to rule * Use sender, statemanager, evalfactory directly * evalApplied and stopApplied * use schedulableAlertRules behind interface * loaded metrics reader * 3 relevant config options * Drop unused scheduler parameter * Rename ruleRoutine to run * Update READMED * Handle long parameter lists * remove dead branch	2024-03-06 13:44:53 -06:00
Alexander Weaver	1bb38e8f95	Alerting: Move ruleRoutine to be a method on ruleInfo (#83866 ) * Move ruleRoutine to ruleInfo file * Move tests as well * swap ruleInfo and scheduler parameters on ruleRoutine * Fix linter complaint, receiver name	2024-03-04 17:15:55 -06:00
Alexander Weaver	f2a9d0a89d	Alerting: Refactor ruleRoutine to take an entire ruleInfo instance (#83858 ) * Make stop a real method * ruleRoutine takes a ruleInfo reference directly rather than pieces of it * Fix whitespace	2024-03-04 15:15:01 -06:00
Matthew Jacobson	2e8c514cfd	Alerting: Stop persisting user-defined templates to disk (#83456 ) Updates Grafana Alertmanager to work with new interface from grafana/alerting#161. This change stops passing user-defined templates to the Grafana Alertmanager by persisting them to disk and instead passes them by string.	2024-03-04 20:12:49 +02:00
Alexander Weaver	fa51724bc6	Alerting: Move alertRuleInfo and tests to new files (#83854 ) Move ruleinfo and tests to new files	2024-03-04 11:24:49 -06:00
Ryan McKinley	3036b50df3	Expressions: expose ConvertDataFramesToResults (#83805 )	2024-03-04 18:22:56 +02:00
Ryan McKinley	5f6bf93dd5	Expressions: Use enumerations rather than strings (#83741 )	2024-03-01 19:38:32 +02:00
Santiago	8ad367e4ad	Chore: Remove redundant error check (#83769 )	2024-03-01 13:28:08 -03:00
Alexander Weaver	a862a4264d	Alerting: Export rule validation logic and make it portable (#83555 ) * ValidateInterval doesn't need the entire config * Validation no longer depends on entire folder now that we've dropped foldertitle from api * Don't depend on entire config struct * Export validate group	2024-02-28 14:40:13 -06:00
Joe Blubaugh	b905777ba9	Alerting: Support deleting rule groups in the provisioning API (#83514 ) * Alerting: feat: support deleting rule groups in the provisioning API Adds support for DELETE to the provisioning API's alert rule groups route, which allows deleting the rule group with a single API call. Previously, groups were deleted by deleting rules one-by-one. Fixes #81860 This change doesn't add any new paths to the API, only new methods. --------- Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-02-28 10:19:02 -05:00
김은빈	96dfb385ca	Grafana: Replace magic number with a constant variable in response status (#80132 ) * Chore: Replace response status with const var * Apply suggestions from code review Co-authored-by: Sofia Papagiannaki <1632407+papagian@users.noreply.github.com> * Add net/http import --------- Co-authored-by: Sofia Papagiannaki <1632407+papagian@users.noreply.github.com>	2024-02-27 18:39:51 +02:00
Takashi Idobe	f209a5a8b8	fix typos (#83414 )	2024-02-26 10:52:44 -07:00
Serge Zaitsev	d0679f0993	Chore: Add support bundle for folders (#83360 ) * add support bundle for folders * fix ProvideService in tests * add a test for collector	2024-02-26 11:27:22 +01:00
George Robinson	a0353b237a	Alerting: Update swagger specs (#83260 )	2024-02-23 11:25:43 +00:00
George Robinson	a564c8c439	Alerting: Keep order of time and mute time intervals consistent (#83257 )	2024-02-22 16:57:20 +00:00
George Robinson	1ed1242358	Alerting: Basic support for time_intervals (#83216 ) This commit adds basic support for time_intervals, as mute_time_intervals is deprecated in Alertmanager and scheduled to be removed before 1.0. It does not add support for time_intervals in API or file provisioning, nor does it support exporting time intervals. This will be added in later commits to keep the changes as simple as possible.	2024-02-22 15:58:56 +00:00
Matthew Jacobson	87ab98ea95	Alerting: Fix panic in provisioning filter contacts by unknown name (#83070 )	2024-02-19 17:30:13 +02:00
Matthew Jacobson	46a77c0074	Alerting: Validate upgraded receivers early to display in preview (#82956 ) Previously receivers were only validated before saving the alertmanager configuration. This is a suboptimal experience for those upgrading with preview as the failed channel upgrade will return an API error instead of being summarized in the table.	2024-02-16 15:17:07 -05:00
William Wernert	fabaff9a24	Alerting: Create metric for rules using simple notifications (#82904 ) --------- Co-authored-by: Matthew Jacobson <matthew.jacobson@grafana.com>	2024-02-16 19:01:49 +02:00
Matthew Jacobson	dfaf6d1e2e	Alerting: Dry-run legacy upgrade on startup (#82835 ) Adds a feature flag (alertingUpgradeDryrunOnStart) that will dry-run the legacy alert upgrade on startup. It is enabled by default. When on legacy alerting, this feature flag will log the results of the legacy alerting upgrade on startup and draw attention to anything in the current legacy alerting configuration that will cause issues when the upgrade is eventually performed. It acts as a log warning for those where action is required before upgrading to Grafana v11 where legacy alerting will be removed.	2024-02-16 11:29:54 -05:00
Matthew Jacobson	e7c6e9c5c9	Alerting: Fix migration edge-case race condition for silences (#81206 ) If the db already has an entry in the kvstore for the silences of an alertmanager before the migration has taken place, then it's possible that the active alertmanager will overwrite the silence file created by the migration before it has a chance to load it into memory. This should not happen normally but is possible in edge-cases. This change opts to bypass the unnecessary step of writing the silences to disk during the migration and instead write them directly to the kvstore. This avoids the race condition entirely and is more correct as we treat the database as the source of truth for AM state.	2024-02-16 10:47:34 -05:00
Gabriel MABILLE	846eadff63	RBAC Search: Replace `userLogin` filter by `namespacedID` filter (#81810 ) * Add namespace ID * Refactor and add tests * Rename maxOneOption -> atMostOneOption * Add ToDo * Remove UserLogin & UserID for NamespaceID Co-authored-by: jguer <joao.guerreiro@grafana.com> * Remove unecessary import of the userSvc * Update pkg/services/accesscontrol/acimpl/service.go * fix 1 -> userID * Update pkg/services/accesscontrol/accesscontrol.go --------- Co-authored-by: jguer <joao.guerreiro@grafana.com>	2024-02-16 11:42:36 +01:00
Sven Kirschbaum	86c618a6d6	Alerting: Escape namespace and group path parameters (#80504 ) Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com>	2024-02-16 09:43:47 +01:00
Matthew Jacobson	118e4a50b7	Alerting: Remove start page of legacy upgrade preview (#82010 ) Alerting: Remove start page of upgrade preview Alerting upgrade page will now always show the summary table even before upgrading any alerts or notification channels. There a few reasons for this: - The information on the start page is redundant as it's now contained in the documentation. - Previously, if some unexpected issue prevented performing a full upgrade, a user would have limited to no means to using the preview tool to help fix the problem. This is because you could not see the summary table until the full upgrade was performed at least once. Now, you can upgrade individual alerts and notification channels from the beginning.	2024-02-15 17:34:00 -05:00
Julien Duchesne	ba63e62311	Alerting: Return provenance of notification templates (#82274 )	2024-02-15 14:35:54 -05:00
William Wernert	b7bbc5058f	Alerting: Don't validate rules on group update if they've only been reordered (#81841 ) --------- Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-02-15 12:03:28 -05:00
Yuri Tseretyan	1eebd2a4de	Alerting: Support for simplified notification settings in rule API (#81011 ) * Add notification settings to storage\domain and API models. Settings are a slice to workaround XORM mapping * Support validation of notification settings when rules are updated * Implement route generator for Alertmanager configuration. That fetches all notification settings. * Update multi-tenant Alertmanager to run the generator before applying the configuration. * Add notification settings labels to state calculation * update the Multi-tenant Alertmanager to provide validation for notification settings * update GET API so only admins can see auto-gen	2024-02-15 09:45:10 -05:00
Alexander Weaver	d4ae10ecc6	Alerting: Small refactor, move unrelated functions out of fetcher (#82459 ) Move unrelated functions out of fetcher	2024-02-14 20:01:32 +02:00
Diego Augusto Molina	ff08c0a790	Chore: improve test readability in ngalert/schedule (#82453 ) Chore: improve test readability	2024-02-14 14:53:32 -03:00
Diego Augusto Molina	9c29e1a783	Alerting: Fix data races and improve testing (#81994 ) * Alerting: fix race condition in (ngalert/sender.ExternalAlertmanager).Run Chore: Fix data races when accessing members of ngalert/state.FakeInstanceStore Chore: Fix data races in tests in ngalert/schedule and enable some parallel tests * Chore: fix linters * Chore: add TODO comment to remove loopvar once we move to Go 1.22	2024-02-14 12:45:39 -03:00
Alexander Weaver	ccb4533a86	Alerting: Remove unused AlertRuleVersionWithPauseStatus (#82383 ) Remove unused AlertRuleVersionWithPauseStatus	2024-02-13 10:56:24 -06:00
Alexander Weaver	99fa064576	Alerting: Emit warning when creating or updating unusually large groups (#82279 ) * Add config for limit of rules per rule group * Warn when editing big groups through normal API * Warn on prov api writes for groups * Wire up comp root, tests * Also add warning to state manager warm * Drop unnecessary conversion	2024-02-13 08:29:03 -06:00
Ryan McKinley	0c6e409350	Chore: Update arrow and prometheus dependencies (#82215 ) * update arrow and prometheus * keep codeowner * use compare * use grafana-plugin-sdk-go v0.210.0 --------- Co-authored-by: ismail simsek <ismailsimsek09@gmail.com>	2024-02-13 01:50:25 +01:00
Karl Persson	1315c67c8b	Team/User: UID migrations (#82298 ) * Add user uid migration to run on every startup to protect against empty values in a upgrade downgrade scenario * Add team uid migration to run on every startup to protect against empty values in a upgrade downgrade scenario * Run team uid migration	2024-02-12 14:48:29 +01:00
Alexander Weaver	5bbe9c6e61	Alerting: Enable group-level rule evaluation jittering by default, remove feature toggle (#82212 ) * remove jitter feature flag * Add an out so users can manually disable jitter * Pass in cfg * Add TODO to remove knob in future	2024-02-09 15:53:58 -06:00
Dan Cech	790e1feb93	Chore: Update test database initialization (#81673 ) * streamline initialization of test databases, support on-disk sqlite test db * clean up test databases * introduce testsuite helper * use testsuite everywhere we use a test db * update documentation * improve error handling * disable entity integration test until we can figure out locking error	2024-02-09 09:35:39 -05:00
Jean-Philippe Quéméner	4dc1ebbb66	fix(alerting): add a proper compare func for location in mute timings (#82153 )	2024-02-08 13:36:09 +01:00
George Robinson	90a26e18db	Alerting: Update Alertmanager to e82436c (#82145 ) This commit updates Alertmanager to commit e82436c, which is based on commit f69a508 from Prometheus Alertmanager.	2024-02-08 11:25:27 +00:00
Matthew Jacobson	dd0ca1263b	Alerting: Include rule uid, title, namespace in unique constraint errors (#82011 ) * Alerting: Include rule_uid, title, namespace_uid in unique constraint errors	2024-02-07 12:55:48 -05:00
Alexander Weaver	843c477899	Alerting: Add exported API to scheduler to access currently loaded rules (#82031 ) * Add exported API to fetch rule definitions from scheduler * Add comment	2024-02-07 09:31:22 -06:00
Yuri Tseretyan	47546a4c72	Alerting: Update API to use folders' full paths (#81214 ) * update GetUserVisibleNamespaces to use FolderSeriver * update GetNamespaceByUID to use FolderService.GetFolders * update GetAlertRulesForScheduling to use FolderService.GetFolders * Update API and GetAlertRulesForScheduling to use the folder's full path * get full path of folder in RouteTestGrafanaRuleConfig * fix escaping of titles for MySQL	2024-02-06 17:12:13 -05:00
Gokhan	cf601fab09	Alerting: Enable sending notifications to a specific topic on Telegram (#79546 ) Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2024-02-06 17:19:22 +02:00
William Wernert	2ea82af6e7	Alerting: Pass in receiver service to API struct (#81978 )	2024-02-06 16:49:47 +02:00
George Robinson	c8ccc4649c	Alerting: Support UTF-8 (#81512 ) This pull request updates our fork of Alertmanager to commit 65bdab0, which is based on commit 5658f8c in Prometheus Alertmanager. It applies the changes from grafana/alerting#155 which removes the overrides for validation of alerts, labels and silences that we had put in place to allow alerts and silences to work for non-Prometheus datasources. However, as this is now supported in Alertmanager with the UTF-8 work, we can use the new upstream functions and remove these overrides. The compat package is a package in Alertmanager that takes care of backwards compatibility when parsing matchers, validating alerts, labels and silences. It has three modes: classic mode, UTF-8 strict mode, fallback mode. These modes are controlled via compat.InitFromFlags. Grafana initializes the compat package without any feature flags, which is the equivalent of fallback mode. Classic and UTF-8 strict mode are used in Mimir. While Grafana Managed Alerts have no need for fallback mode, Grafana can still be used as an interface to manage the configurations of Mimir Alertmanagers and view configurations of Prometheus Alertmanager, and those installations might not have migrated or being running on older versions. Such installations behave as if in classic mode, and Grafana must be able to parse their configurations to interact with them for some period of time. As such, Grafana uses fallback mode until we are ready to drop support for outdated installations of Mimir and the Prometheus Alertmanager.	2024-02-06 08:33:47 +00:00
William Wernert	2ab7d3c725	Alerting: Receivers API (read only endpoints) (#81751 ) * Add single receiver method * Add receiver permissions * Add single/multi GET endpoints for receivers * Remove stable tag from time intervals See end of PR description here: https://github.com/grafana/grafana/pull/81672	2024-02-05 20:12:15 +02:00
Ryan McKinley	9c9e5e68c8	User: Add uid colum to user table (#81615 )	2024-02-01 18:14:10 -08:00
Yuri Tseretyan	d1073deefd	Alerting: Time intervals API (read only endpoints) (#81672 ) * declare new API and models GettableTimeIntervals, PostableTimeIntervals * add new actions alert.notifications.time-intervals:read and alert.notifications.time-intervals:write. * update existing alerting roles with the read action. Add to all alerting roles. * add integration tests	2024-02-01 15:17:13 -05:00
William Wernert	7e939401dc	Alerting: Introduce initial common receiver service (#81211 ) * Create locking config store that mimics existing provisioning store * Rename existing receivers(_test).go * Introduce shared receiver group service * Fix test * Move query model to models package * ReceiverGroup -> Receiver * Remove locking config store * Move convert methods to compat.go * Cleanup	2024-02-01 14:42:59 -05:00
George Robinson	0726c7c3fa	Alerting: Prevent inhibition rules in Grafana Alertmanager (#81712 ) This commit prevents saving configurations containing inhibition rules in Grafana Alertmanager. It does not reject inhibition rules when using external Alertmanagers, such as Mimir. This meant the validation had to be put in the MultiOrgAlertmanager instead of in the validation of PostableUserConfig. We can remove this when inhibition rules are supported in Grafana Managed Alerts.	2024-02-01 14:53:15 +00:00
Matthew Jacobson	0ce1ccd6f9	Alerting: Fix inconsistent AM raw config when applied via sync vs API (#81655 ) AM config applied via API would use the PostableUserConfig as the AM raw config and also the hash used to decide when the AM config has changed. However, when applied via the periodic sync the PostableApiAlertingConfig would be used instead. This leads to two issues: - Inconsistent hash comparisons when modifying the AM causing redundant applies. - GetStatus assumed the raw config was PostableUserConfig causing the endpoint to return correctly after a new config is applied via API and then nothing once the periodic sync runs. Note: Technically, the upstream GrafanaAlertamanger GetStatus shouldn't be returning PostableUserConfig or PostableApiAlertingConfig, but instead GettableStatus. However, this issue required changes elsewhere and is out of scope.	2024-01-31 21:05:30 +02:00
Ashley Harrison	39057552dc	QueryField: Handle autocomplete better (#81484 ) * extract out function + add unit tests * add feature toggle and default it to on	2024-01-31 10:01:20 +00:00
Yuri Tseretyan	131c72d655	Alerting: Fix scheduler to group folders by the unique key (orgID and UID) (#81303 )	2024-01-30 17:14:11 -05:00
Sofia Papagiannaki	89d3b55bec	Folders: Reduce DB queries when counting and deleting resources under folders (#81153 ) * Add folder store method for fetching all folder descendants * Modify GetDescendantCounts() to fetch folder descendants at once * Reduce DB calls when counting library panels under dashboard * Reduce DB calls when counting dashboards under folder * Reduce DB calls during folder delete * Modify folder registry to count/delete entities under multiple folders * Reduce DB calls when counting * Reduce DB calls when deleting	2024-01-30 18:26:34 +02:00
William Wernert	de662810cf	Alerting: Create instance of alert rule generator in historian annotation tests (#81394 ) * Create generator variable to ensure closures have correct context	2024-01-29 11:22:43 -05:00
idafurjes	f44592a97a	Remove folderID from service tests (#80615 ) * Remove folderID from service tests * Remove folderID from ngalert migration tests * Remove tests related to folderIDs * Roll back change Before removing FolderID from this test, we need to adjust the code * Remove FolderID from publicdashboard pkg * Add back annotations test	2024-01-26 17:36:35 +02:00
Gabriel MABILLE	722b78f3e0	RBAC: Add userLogin filter to the permission search endpoint (#81137 ) * RBAC: Search add user login filter * Switch to a userService resolving instead * Remove unused error * Fallback to use the cache * account for userID filter * Account for the error * snake case * Add test cases * Add api tests * Fix return on error * Re-order imports	2024-01-26 09:43:16 +01:00
Sofia Papagiannaki	b1eec36df3	Alerting: Fix authorisation to use namespace UIDs for scope (#81231 )	2024-01-25 15:19:51 -05:00
idafurjes	7e5544ab21	Add MFolderIDsServiceCount to count folderIDs in services pkg (#81237 )	2024-01-25 11:10:35 +01:00
Sofia Papagiannaki	478d7d58fa	Nested folders: Allow creating folders with duplicate names in different locations (#77076 ) * Add API test * Add move tests * Fix create folder * Fix move * Fix test * Drop and re-create index so that allows a folder to contain a dashboard and a subfolder with same name * Get folder by title defaults to root folder and optionally fetches folder by provided parent folder * Apply suggestions from code review	2024-01-25 11:29:56 +02:00
William Wernert	2203bc2a3d	Alerting: Refactor provisioning tests/fakes (#81205 ) * Fix up test Alertmanager config JSON * Move fake AM config and provisioning stores to fakes package	2024-01-24 17:15:55 -05:00
Matthew Jacobson	71e70c424f	Alerting: During legacy migration reduce the number of created silences (#78505 ) * Alerting: During legacy migration reduce the number of created silences During legacy migration every migrated rule was given a label rule_uid=<uid>. This was used to silence DatasourceError/DatasourceNoData alerts for migrated rules that had either ExecutionErrorState/NoDataState set to keep_state, respectively. This could potentially create a large amount of silences and a high cardinality label. Both of these scenarios have poor outcomes for CPU load and latency in unified alerting. Instead, this change creates one label per ExecutionErrorState/NoDataState when they are set to keep_state as well as two silence rules, if rules with said labels were created during migration. These silence rules are: - __legacy_silence_error_keep_state__ = true - __legacy_silence_nodata_keep_state__ = true This will drastically reduce the number of created silence rules in most cases as well as not create the potentially high cardinality label `rule_uid`.	2024-01-24 15:56:19 -05:00
Santiago	fbbda6c05e	Alerting: Retry readiness check to the remote Alertmanager on 5xx status code responses (#81174 )	2024-01-24 21:39:06 +01:00
George Robinson	05d858635c	Alerting: Add metric for inhibition rules (#81119 ) This commit adds a metric for the number of inhibition rules. It matches the metric added upstream in #3681.	2024-01-23 19:43:17 +00:00
Jean-Philippe Quéméner	aa25776f81	Alerting: Add a feature flag to periodically save states (#80987 )	2024-01-23 17:03:30 +01:00
George Robinson	85b9edcd28	Alerting: Fix incorrect initialization of logger (#81099 )	2024-01-23 17:29:38 +02:00
Marcus Efraimsson	6768c6c059	Chore: Remove public vars in setting package (#81018 ) Removes the public variable setting.SecretKey plus some other ones. Introduces some new functions for creating setting.Cfg.	2024-01-23 12:36:22 +01:00
Jean-Philippe Quéméner	eb7e1216a1	feat(alerting): add async state persister (#80763 )	2024-01-22 13:07:11 +01:00

... 2 3 4 5 6 ...

1612 Commits