grafana

mirror of https://github.com/grafana/grafana.git synced 2024-11-26 02:40:26 -06:00

Author	SHA1	Message	Date
Kyle Brandt	2f13c851e4	SSE: (Chore/Instrumentation) Add ds_queries_total metric and move met… (#66695 ) * SSE: (Chore/Instrumentation) Add ds_queries_total metric and move metrics to service	2023-04-17 16:12:44 -07:00
George Robinson	19ebb079ba	Alerting: Add limits and filters to Prometheus Rules API (#66627 ) This commit adds support for limits and filters to the Prometheus Rules API. Limits: It adds a number of limits to the Grafana flavour of the Prometheus Rules API: - `limit` limits the maximum number of Rule Groups returned - `limit_rules` limits the maximum number of rules per Rule Group - `limit_alerts` limits the maximum number of alerts per rule It sorts Rule Groups and rules within Rule Groups such that data in the response is stable across requests. It also returns summaries (totals) for all Rule Groups, individual Rule Groups and rules. Filters: Alerts can be filtered by state with the `state` query string. An example of an HTTP request asking for just firing alerts might be `/api/prometheus/grafana/api/v1/rules?state=alerting`. A request can filter by two or more states by adding additional `state` query strings to the URL. For example `?state=alerting&state=normal`. Like the alert list panel, the `firing`, `pending` and `normal` state are first compared against the state of each alert rule. All other states are ignored. If the alert rule matches then its alert instances are filtered against states once more. Alerts can also be filtered by labels using the `matcher` query string. Like `state`, multiple matchers can be provided by adding additional `matcher` query strings to the URL. The match expression should be parsed using existing regular expression and sent to the API as URL-encoded JSON in the format: { "name": "test", "value": "value1", "isRegex": false, "isEqual": true } The `isRegex` and `isEqual` options work as follows: \| IsEqual \| IsRegex \| Operator \| \| ------- \| -------- \| -------- \| \| true \| false \| = \| \| true \| true \| =~ \| \| false \| true \| !~ \| \| false \| false \| != \|	2023-04-17 17:45:06 +01:00
Arati R	cab3ba519a	NestedFolders: Add folder service registry with dashboard service implementation (#65033 ) * Delete folders, dashboards with registry service Co-authored-by: Serge Zaitsev <hello@zserge.com> * Update signature of ProvideDashboardServiceImpl * Regenerate mockery file * Add test for DeleteInFolder * Add test for DeleteDashboardsInFolder * Delete child dashboard associations via registry * Add validation of folder uid and org id --------- Co-authored-by: Serge Zaitsev <hello@zserge.com>	2023-04-14 11:17:23 +02:00
Yuri Tseretyan	afd52d0866	Alerting: use alerting GrafanaReceiver and BuildReceiverConfiguration in Grafana (#65224 ) * replace receiver errors with one from alerting * add the converter to alerting models * update buildReceiverIntegration to accept GrafanaReceiver --------- Co-authored-by: George Robinson <george.robinson@grafana.com>	2023-04-13 12:25:32 -04:00
gotjosh	2bbf0c9de4	Alerting: Allow Rules to Schedule to be filtered by Rule Group (#59990 ) * Alerting: Allow Rules to Schedule to be filtered by Rule Group	2023-04-13 12:55:42 +01:00
Michael Mandrus	5626461b3c	Caching: Refactor enterprise query caching middleware to a wire service (#65616 ) * define initial service and add to wire * update caching service interface * add skipQueryCache header handler and update metrics query function to use it * add caching service as a dependency to query service * working caching impl * propagate cache status to frontend in response * beginning of improvements suggested by Lean - separate caching logic from query logic. * more changes to simplify query function * Decided to revert renaming of function * Remove error status from cache request * add extra documentation * Move query caching duration metric to query package * add a little bit of documentation * wip: convert resource caching * Change return type of query service QueryData to a QueryDataResponse with Headers * update codeowners * change X-Cache value to const * use resource caching in endpoint handlers * write resource headers to response even if it's not a cache hit * fix panic caused by lack of nil check * update unit test * remove NONE header - shouldn't show up in OSS * Convert everything to use the plugin middleware * revert a few more things * clean up unused vars * start reverting resource caching, start to implement in plugin middleware * revert more, fix typo * Update caching interfaces - resource caching now has a separate cache method * continue wiring up new resource caching conventions - still in progress * add more safety to implementation * remove some unused objects * remove some code that I left in by accident * add some comments, fix codeowners, fix duplicate registration * fix source of panic in resource middleware * Update client decorator test to provide an empty response object * create tests for caching middleware * fix unit test * Update pkg/services/caching/service.go Co-authored-by: Arati R. <33031346+suntala@users.noreply.github.com> * improve error message in error log * quick docs update * Remove use of mockery. Update return signature to return an explicit hit/miss bool * create unit test for empty request context * rename caching metrics to make it clear they pertain to caching * Update pkg/services/pluginsintegration/clientmiddleware/caching_middleware.go Co-authored-by: Marcus Efraimsson <marcus.efraimsson@gmail.com> * Add clarifying comments to cache skip middleware func * Add comment pointing to the resource cache update call * fix unit tests (missing dependency) * try to fix mystery syntax error * fix a panic * Caching: Introduce feature toggle to caching service refactor (#66323) * introduce new feature toggle * hide calls to new service behind a feature flag * remove licensing flag from toggle (misunderstood what it was for) * fix unit tests * rerun toggle gen --------- Co-authored-by: Arati R. <33031346+suntala@users.noreply.github.com> Co-authored-by: Marcus Efraimsson <marcus.efraimsson@gmail.com>	2023-04-12 12:30:33 -04:00
Kyle Brandt	e78be44e1a	SSE: Dataplane Compliance (#65927 ) Takes a specific code path for data that identifies itself as dataplane instead of "guessing" what the data is. The data must identify itself by being in the dataplane by having both the following frame metadata properties: - TypeVersion property that is greater than 0.0 - 'Type' property The flag is disableSSEDataplane and disables this functionality and uses the old code for all queries regardless. See https://github.com/grafana/grafana-plugin-sdk-go/blob/main/data/contract_docs/contract.md for dataplane details.	2023-04-12 12:24:34 -04:00
Matthew Jacobson	63187fae0c	Alerting: Remove and revert flag alertingBigTransactions (#65976 ) * Alerting: Remove and revert flag alertingBigTransactions This is a partial revert of #56575 and a removal of the `alertingBigTransactions` flag. Real-word use has seen no clear performance incentive to maintain this flag. Lowered db connection count came at the cost of significant increase in CPU usage and query latency. * Fix lint backend * Removed last bits of alertingBigTransactions --------- Co-authored-by: Armand Grillet <2117580+armandgrillet@users.noreply.github.com>	2023-04-06 18:06:25 +02:00
gotjosh	1c3ce0735f	Alerting: Tiny refactor on the eval and schedule packages (#66130 ) * Alerting: Tiny refactor on the eval and schedule packages two very small things: - We had a constructor on something called a `Context` which is not a `context.Context` so let's just name that constructor `NewContext` - The user that we use to run query evaluations is the same (with some variation) abstract it to a function so that it can be re-used when necessary. * Update pkg/services/ngalert/schedule/schedule.go Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com> * Update pkg/services/ngalert/schedule/schedule.go Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com> --------- Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com>	2023-04-06 16:02:28 +01:00
Matthew Jacobson	85f738cdf9	Alerting: Add endpoint to revert to a previous alertmanager configuration (#65751 ) * Alerting: Add endpoint to revert to a previous alertmanager configuration This endpoint is meant to be used in conjunction with /api/alertmanager/grafana/config/history to revert to a previously applied alertmanager configuration. This is done by ID instead of raw config string in order to avoid secure field complications.	2023-04-05 14:10:03 -04:00
Alexander Weaver	fb520edd72	Alerting: Use a completely isolated context for state history writes (#64989 ) * Add fresh context with timeout and same log properties, re-derive logger * Unify timeout constants * Move ctx after shortcut that got added through rebasing * Unify timeouts * Port opentracing's SpanFromContext and ContextFromSpan to the grafana tracing package * Support both opentracing and otel variants * Better document why we're creating a new ctx * Add new func to FakeSpan which was added after rebase * Support grafana-specific traceID key in both tracer implementations	2023-04-04 16:41:46 -05:00
George Robinson	bd29071a0d	Revert "Alerting: Add limits to the Prometheus Rules API" (#65842 )	2023-04-03 15:20:37 +00:00
George Robinson	d96b0a71d3	Alerting: Add limits to the Prometheus Rules API (#65169 ) This commit adds a number of limits to the Grafana flavor of the Prometheus Rules API: 1. `limit` limits the maximum number of Rule Groups returned 2. `limit_rules` limits the maximum number of rules per Rule Group 3. `limit_alerts` limits the maximum number of alerts per rule It sorts Rule Groups and rules within Rule Groups such that data in the response is stable across requests. It also returns summaries (totals) for all Rule Groups, individual Rule Groups and rules.	2023-04-03 10:17:02 +01:00
Santiago	aba91d3053	Alerting: Fetch all applied alerting configurations (#65728 ) * WIP * skip invalid historic configurations instead of erroring * add warning log when bad historic config is found * remove unused custom marshaller for GettableHistoricUserConfig * add id to historic user config, move limit check to store, fix typo * swagger spec	2023-03-31 17:43:04 -03:00
Alexander Weaver	da4832724e	Alerting: Delete stub for SQL alert state history backend (#65667 ) Delete stub for SQL backend	2023-03-31 11:15:56 -05:00
Matthew Jacobson	b9dc04139a	Alerting: Respect "For" Duration for NoData alerts (#65574 ) * Alerting: Respect "For" Duration for NoData alerts This change modifies `resultNoData` to be more inline with the logic of the other state handlers. The main effects of this are: 1) NoData states with NoDataState config set to Alerting will respect "For" duration. 2) Prevents zero value in StartsAt and EndsAt for alerts that have only even been in normal state. This includes state transitions from NoDataState=OK and ExecErrState=OK. 3) Better state transition logging.	2023-03-31 19:05:15 +03:00
Steve Simpson	04336d53a9	Alerting: Update prometheus version (#65688 )	2023-03-31 16:34:35 +02:00
Yuri Tseretyan	622c23716a	Alerting: Use logger with context in the state cache (#65663 )	2023-03-31 10:11:30 -04:00
Alexander Weaver	b2abb63286	Alerting: Introduce proper feature toggles for common state history backend combinations (#65497 ) * define 3 feature toggles for rollout phases * Pass feature toggles along * Implement first feature toggle * Try a different strategy with fall-throughs to specific configurations * Apply toggle overrides once outside of backend composition * Emit log messages when we coerce backends * Run code generator for feature toggle files * Improve wording in flag descs * Re-run generator * Use code-generated constants instead of plain strings * Use converted enum values rather than strings for pre-parsing	2023-03-30 13:53:21 -05:00
Alexander Weaver	5e87ea745d	Alerting: Fix and re-enable `filters instance labels in log line` test (#65618 ) Fix and reenable test	2023-03-30 09:02:18 -05:00
Dimitris Sotirakis	e758b017d0	Alerting: Disable `filters instance labels in log line` test (#65610 ) * Disable filters instance labels in log line test * Add drone reference	2023-03-30 16:04:29 +03:00
Yuri Tseretyan	9eaffdf5a8	Alerting: Remove dependency on alerting package in definitions (#65390 ) * move export rules to definitions package * move provisioning contact point methods to provisioning package * move AlertRuleGroupWithFolderTitle to ngalert models and adapter functions to api's compat * move rule_types files back to where they were before.	2023-03-29 13:34:59 -04:00
Alexander Weaver	a416100abc	Alerting: No longer index state history log streams by instance labels (#65474 ) * Remove private labels * No longer index by instance labels * Labels are now invariant, only build them once * Remove bucketing since everything is in a single stream * Refactor statesToStreams to only return a single unified log stream * Don't query on labels that no longer exist * Move selector logic to loki layer, genericize client to work in terms of straight logQL * Add support for line-level label filters in query * Combine existing selector tests for better parallelism * Tests for logQL construction * Underscore instead of dot for unwrapping labels in logql	2023-03-29 11:52:11 -05:00
Santiago	7b92849508	Alerting: Add CustomDetails field in PagerDuty contact point (#64860 ) * Alerting: Add CustomDetails for PagerDuty * fix default value for 'severity' from 'error' to 'critical' * minimal docs for notifiers, specifying config for PagerDuty * replace notifier -> integration * replace notifier -> integration	2023-03-29 10:35:01 -03:00
Alexander Weaver	de1637afe5	Alerting: Add alert instance labels to Loki log lines in addition to stream labels (#65403 ) Add instance labels to log line	2023-03-28 08:57:51 -05:00
Alexander Weaver	dd04757fc9	Alerting: Add "backend" label to state history writes metrics (#65395 ) * Add backend label to state history writes metrics * Update test expectations	2023-03-28 08:49:51 -05:00
Giuseppe Guerra	a89202eab2	Plugins: Improve instrumentation by adding metrics and tracing (#61035 ) * WIP: Plugins tracing * Trace ID middleware * Add prometheus metrics and tracing to plugins updater * Add TODOs * Add instrumented http client * Add tracing to grafana update checker * Goimports * Moved plugins tracing to middleware * goimports, fix tests * Removed X-Trace-Id header * Fix comment in NewTracingHeaderMiddleware * Add metrics to instrumented http client * Add instrumented http client options * Removed unused function * Switch to contextual logger * Refactoring, fix tests * Moved InstrumentedHTTPClient and PrometheusMetrics to their own package * Tracing middleware: handle errors * Report span status codes when recording errors * Add tests for tracing middleware * Moved fakeSpan and fakeTracer to pkg/infra/tracing * Add TestHTTPClientTracing * Lint * Changes after PR review * Tests: Made "ended" in FakeSpan private, allow calling End only once * Testing: panic in FakeSpan if span already ended * Refactoring: Simplify Grafana updater checks * Refactoring: Simplify plugins updater error checks and logs * Fix wrong call to checkForUpdates -> instrumentedCheckForUpdates * Tests: Fix wrong call to checkForUpdates -> instrumentedCheckForUpdates * Log update checks duration, use Info log level for check succeeded logs * Add plugin context span attributes in tracing_middleware * Refactor prometheus metrics as httpclient middleware * Fix call to ProvidePluginsService in plugins_test.go * Propagate context to update checker outgoing http requests * Plugin client tracing middleware: Removed operation name in status * Fix tests * Goimports tracing_middleware.go * Goimports * Fix imports * Changed span name to plugins client middleware * Add span name assertion in TestTracingMiddleware * Removed Prometheus metrics middleware from grafana and plugins updatechecker * Add span attributes for ds name, type, uid, panel and dashboard ids * Fix http header reading in tracing middlewares * Use contexthandler.FromContext, add X-Query-Group-Id * Add test for RunStream * Fix imports * Changes from PR review * TestTracingMiddleware: Changed assert to require for didPanic assertion * Lint * Fix imports	2023-03-28 11:01:06 +02:00
Serge Zaitsev	0beb768427	Chore: Remove result fields from ngalert (#65410 ) * remove result fields from ngalert * remove duplicate imports	2023-03-28 10:34:35 +02:00
Yuri Tseretyan	ec4152c7e5	Alerting: Remove dependency on secrets in definitions package (#65391 )	2023-03-27 16:35:54 -04:00
Yuri Tseretyan	52a0f59706	Alerting: introduce AlertQuery in definitions package (#63825 ) * copy AlertQuery from ngmodels to the definition package * replaces usages of ngmodels.AlertQuery in API models * create a converter between models of AlertQuery --------- Co-authored-by: Alex Moreno <alexander.moreno@grafana.com>	2023-03-27 11:55:13 -04:00
Alexander Weaver	07368dec74	Alerting: Fix attachment of external labels to Loki state history log streams (#65140 ) Fix attachment of external labels, add tests	2023-03-21 18:00:59 -05:00
Alexander Weaver	bf54f2672e	Alerting: Switch to snappy-compressed-protobuf for outgoing push requests to Loki (#65077 ) * Encode with snappy, always * JSON encoder type * Headers * Copy labels formatter from promtail * Implement snappy-proto encoding * Create encoder interface, test both encoders, choose snappy-proto by default * Make encoder configurable at the LokiCfg level * Export both encoders * Touch up comment and tests * Drop unnecessary conversions after move to plain strings to appease linter	2023-03-21 13:38:42 -05:00
Alexander Weaver	cc7e5ce62e	Alerting: Fix ambiguous handling of equals in labels when bucketing Loki state history streams (#65013 ) * Use JSON instead of data.Labels string format as label repr * Drop debug log line	2023-03-21 12:33:27 -05:00
Alexander Weaver	e39d7f44c9	Alerting: Elide requests to Loki if nothing should be recorded (#65011 ) Exit early if no log streams or annotations	2023-03-21 09:30:56 -05:00
Alexander Weaver	40c5713cbd	Vendor errors.Join from Go standard library to avoid version incompatibilities (#64985 ) Vendor errors.Join from std lib	2023-03-17 14:07:58 -05:00
Alexander Weaver	a31672fa40	Alerting: Create new state history "fanout" backend that dispatches to multiple other backends at once (#64774 ) * Rename RecordStatesAsync to Record * Rename QueryStates to Query * Implement fanout writes * Implement primary queries * Simplify error joining * Add test for query path * Add tests for writes and error propagation * Allow fanout backend to be configured * Touch up log messages and config validation * Consistent documentation for all backend structs * Parse and normalize backend names more consistently against an enum * Touch-ups to documentation * Improve clarity around multi-record blocking * Keep primary and secondaries more distinct * Rename fanout backend to multiple backend * Simplify config keys for multi backend mode	2023-03-17 12:41:18 -05:00
Alexander Weaver	9bcf8819d3	Alerting: Handful of small adjustments to log levels and parameters (#64572 ) Calculate duration earlier in scheduler	2023-03-17 12:15:49 +00:00
gotjosh	02a8f62021	Alerting: Fix stats that display alert count when using unified alerting (#64852 ) * Alerting: Fix stats when using unified alerting	2023-03-17 11:19:18 +00:00
George Robinson	0b506b4ccc	Alerting: Update github.com/grafana/alerting (#64882 )	2023-03-16 13:59:35 +00:00
Yuri Tseretyan	85a954cd81	Alerting: Update scheduler to get updates only from database (#64635 ) * stop using the scheduler's Update and Delete methods all communication must be via the database * update scheduler's registry to calculate diff before re-setting the cache * update fetcher to return the diff generated by registry * update processTick to update rule eval routine if the rule was updated and it is not going to be evaluated at this tick. * remove references to the scheduler from api package * remove unused methods in the scheduler	2023-03-14 18:02:51 -04:00
Alexander Weaver	faef3a8258	Alerting: Log error but don't fail initialization if state history connection test fails (#64699 ) Don't return init error if ping fails, add tests	2023-03-13 15:54:46 -05:00
Andrej Ocenas	6647217208	Phlare: Use enum config to send deduplicated func and filenames (#64435 )	2023-03-13 11:06:04 +01:00
Jean-Philippe Quéméner	fb5ed0b0b3	Alerting: fix flaky cache test (#64499 )	2023-03-09 06:08:05 -05:00
Ryan McKinley	42e7ec9fe4	Chore: cleanup dashboard service names (#64442 )	2023-03-08 14:37:45 -05:00
George Robinson	0c8876c3a2	Alerting: Return errors when expanding templates (#63662 ) This commit changes the state package so that errors encountered while expanding templates for custom labels and annotations are returned from the function. This is not used at present, but will be used in the future as we look at how to offer better feedback to users who don't have access to logs, for example our customers who use Hosted Grafana.	2023-03-08 12:25:02 +00:00
Alexander Weaver	4a1c18abf6	Alerting: Fix intermittency when seeding database in rule store tests (#64322 ) Force unique IDs when seeding database	2023-03-07 09:40:55 -05:00
George Robinson	ed71012ced	Alerting: Fix Classic Conditions $values variable (#64243 ) This commit fixes a bug in the $values variable in notification templates when using Classic Conditions. Since Classic Conditions are not multi-dimensional, the values of each series that exceeded the condition should be available as a RefID and offset. For example, B0, B1, etc. However, this bug meant that instead just a single condition would be printed as B, not B0.	2023-03-06 12:08:00 -05:00
Alexander Weaver	19d01dff91	Alerting: Expose Prometheus metrics for persisting state history (#63157 ) * Create historian metrics and dependency inject * Record counter for total number of state transitions logged * Track write failures * Track current number of active write goroutines * Record histogram of how long it takes to write history data * Don't copy the registerer * Adjust naming of write failures metric * Introduce WritesTotal to complement WritesFailedTotal * Measure TransitionsFailedTotal to complement TransitionsTotal * Rename all to state_history * Remove redundant Total suffix * Increment totals all the time, not just on success * Drop ActiveWriteGoroutines * Drop PersistDuration in favor of WriteDuration * Drop unused gauge * Make writes and writesFailed per org * Add metric indicating backend and a spot for future metadata * Drop _batch_ from names and update help * Add metric for bytes written * Better pairing of total + failure metric updates * Few tweaks to wording and naming * Record info metric during composition * Create fakeRequester and simple happy path test using it * Blocking test for the full historian and test for happy path metrics * Add tests for failure case metrics * Smoke test for full annotation persistence * Create test for metrics on annotation persistence, both happy and failing paths * Address linter complaints * More linter complaints * Remove unnecessary whitespace * Consistency improvements to help texts * Update tests to match new descs	2023-03-06 10:40:37 -06:00
gotjosh	5422f7cf56	Alerting: Add metrics for active receiver and integrations (#64050 ) * Alerting: Add metrics for active receiver and integrations Introduces metrics that allows us to track the number of configured receivers and integration in the Alertmanager for all orgs. As a bonus, I realised that the alert reception metrics where not being exported nor collected. This does that too.	2023-03-06 16:37:07 +00:00
Serge Zaitsev	0bdb105df2	Chore: Remove xorcare/pointer dependency (#63900 ) * Chore: remove pointer dependency * fix type casts * deprecate xorcare/pointer library in linter * rooky mistake	2023-03-06 05:23:15 -05:00
Ieva	a52999a886	Access Control: revert to using folder store from the scope resolvers (#64132 ) * revert to using folder store from the resolvers * fixing tests after revert * api test fixes --------- Co-authored-by: Kristin Laemmert <mildwonkey@users.noreply.github.com>	2023-03-03 10:56:33 -05:00
Yuri Tseretyan	e760f22402	Alerting: Use background context for maintenance function (#64065 )	2023-03-02 14:19:52 -05:00
Emil Tullstedt	10ee900beb	Errors: Remove direct dependencies on github.com/pkg/errors (#64026 ) Co-authored-by: Sofia Papagiannaki <1632407+papagian@users.noreply.github.com>	2023-03-02 16:28:10 +01:00
Kristin Laemmert	bb798e24f3	chore(services): replace dependencies on dashboard store with dashboard service (#63937 ) * chore(services): replace dependencies on dashboard store with dashboard service This continues the backend service/store split by replacing dashboard store dependencies with service dependencies. the folder service remains the single exception for now; otherwise we'd have a dependency cycle between the folder and dashboard services. I have some ideas for that, but I'll take care of all the easy parts first. While doing this, I identified and removed a number of unused arguments from the following functions: NewFolderNameScopeResolver NewFolderIDScopeResolver NewFolderUIDScopeResolver NewDashboardIDScopeResolver NewDashboardUIDScopeResolver resolveDashboardScope I have a small enterprise PR to support this commit. * lingering fmt	2023-03-02 08:09:57 -05:00
Yuri Tseretyan	5e2a661dec	Alerting: update API models to user NoDataState and ExecutionErrorState from definitions instead of models (#63824 )	2023-02-28 16:21:41 -05:00
Yuri Tseretyan	f561e71de8	Alerting: decouple api models from domain\dto models: separate Provenance status + converters (#63594 ) * move conversions of domain models to api models and reverse from definition package to api package	2023-02-27 17:57:15 -05:00
Alexander Weaver	e77621649d	Alerting: Instrument outgoing state history requests using weaveworks/common (#63600 ) * Loki backend and client depend on a requester * Instrument all requests to loki using weaveworks TimedClient * Construct collector in metrics package	2023-02-23 17:52:02 -06:00
Yuri Tseretyan	98e1aeaebd	Alerting: Fix client to external Alertmanager to correctly build URL for Mimir Alertmanager (#63676 )	2023-02-23 13:55:26 -05:00
Emil Tullstedt	3abaf32cf2	Chore: Upgrade golangci-lint to v1.51.2 (#63630 ) Co-authored-by: Sofia Papagiannaki <1632407+papagian@users.noreply.github.com>	2023-02-23 15:10:03 +01:00
Alex Moreno	f60dc4441f	Alerting: Add status label to GroupRules metric (#63454 ) * Add status label to GroupRules metric * Add state (active and paused) label to GrouRules * Add active/paused metrics tests	2023-02-23 12:38:27 +01:00
George Robinson	f93a9c794d	Alerting: Fix incorrect comment in eval.go (#63510 ) This commit fixes an incorrect comment in the Result struct in eval.go that I had written some time ago. The comment now documents the actual behaviour and content of this field.	2023-02-21 15:42:04 +00:00
bla2ej	56c8661929	Alerting: Get alert rules on faults (#61248 ) (#63051 ) * Alerting: get alert rules on faults (#61248) Two functions used to fetch alert rules from DB are updated: - GetAlertRulesForScheduling - ListAlertRules Rows are scanned one by one so good ones are returned. Common Error is logged with indication how many rules failed on deserialization. Resolved: #61248 * updates from review comments	2023-02-21 08:54:20 -06:00
George Robinson	9f2fb3fa27	Alerting: Add filter and remove funcs for custom labels and annotations (#63437 ) This commit adds filterLabels, filterLabelsRe, removeLabels, and removeLabelsRe functions to templates for custom labels and annotations. It allows for use cases such as removing all private labels.	2023-02-20 14:40:26 +00:00
George Robinson	c637a5543e	Alerting: Rename caps to captures as cap is a reserved word (#63432 )	2023-02-20 10:08:36 +00:00
George Robinson	aacf9da969	Alerting: Change Data to use Labels instead of map[string]string (#63431 ) This commit changes the Data struct in template.go to use Labels instead of map[string]string. It changes how labels are printed when using {{ .Labels }} from map[foo:bar bar:baz] to foo=bar, bar=baz.	2023-02-20 10:08:23 +00:00
George Robinson	0a01391ebe	Alerting: Small readability improvements to template.go (#63422 ) * Alerting: Small readability improvements to template.go * Fix lint	2023-02-20 09:24:11 +00:00
George Robinson	0659134793	Alerting: Better printing of labels (#63348 ) This commit changes how labels are printed in templates for custom annotations and labels from map[foo:bar bar:baz] to foo=bar, bar=baz. Labels are comma separated, and sorted in increasing order.	2023-02-16 12:04:15 -05:00
George Robinson	9e86916d48	Alerting: Move templating to template package (#63347 ) This commit moves templating from the state package to a sub-package called template. This sub-package will be the logical package for future ease-of-use improvements to templating custom annotations and labels.	2023-02-16 17:16:36 +01:00
Alexander Weaver	958fb2c50a	Alerting: Unify structs in Loki client and make them more consistent with Prometheus (#63055 ) * Use existing row struct instead of [2]string, add deserialization helper * Replace Stream struct with stream struct which is exactly the same * Drop unused status field * Don't export queryRes and queryData * Tests for custom marshalling * Rename row fields to T and V for consistency with prometheus samples * Rename row to sample	2023-02-11 05:17:44 -06:00
George Robinson	1f984409a2	Alerting: Fix a bug taking screenshots with Dashboard UID (#63220 ) This commit fixes a bug where Grafana would fail to take a screenshot if the same Dashboard UID was present across two or more different orgs.	2023-02-09 15:23:01 -05:00
Steve Simpson	4d1a2c3370	Alerting: Move `rule_groups_rules` metric from State to Scheduler. (#63144 ) The `rule_groups_rules` metric is currently defined and computed by `State`. It makes more sense for this metric to be computed off of the configured rule set, not based on the rule evaluation state. There could be an edge condition where a rule does not have a state yet, and so is uncounted. Additionally, we would like this metric (and others), to have a `rule_group` label, and this is much easier to achieve if the metric is produced from the `Scheduler` package.	2023-02-09 17:05:19 +01:00
suntala	49b3027049	Chore: Remove Result field from datasources (#63048 ) * Remove Result field from AddDataSourceCommand * Remove DatasourcesPermissionFilterQuery Result * Remove GetDataSourceQuery Result * Remove GetDataSourcesByTypeQuery Result * Remove GetDataSourcesQuery Result * Remove GetDefaultDataSourceQuery Result * Remove UpdateDataSourceCommand Result	2023-02-09 15:49:44 +01:00
Alexander Weaver	f80bf11782	Alerting: Make time range query parameters not required when querying Loki (#62985 ) * Make from and to not required * Move default range calculation up to loki.go	2023-02-07 14:26:43 -06:00
Santiago	955c7b13ea	Alerting: Change error log to warning and apply correct format when updating historic config (#62973 ) * change error to warning, apply correct format * Update pkg/services/ngalert/store/alertmanager.go Co-authored-by: George Robinson <george.robinson@grafana.com> * different wording and data --------- Co-authored-by: George Robinson <george.robinson@grafana.com>	2023-02-07 11:04:05 -03:00
Joey Tawadrous	121260e0dd	Parca: Use data query schema (#62840 ) * Parca data query schema * Remove groupBy	2023-02-07 09:56:21 +00:00
Alexander Weaver	0efb84617e	Alerting: Create benchmarking test for state.ProcessEvalResults (#62041 ) * Create benchmark for ProcessEvalResults * Simplify the test	2023-02-03 15:38:08 -06:00
Yuri Tseretyan	f066e8cdcd	Alerting: Update to alerting 20230203015918-0e4e2675d7aa (after refactoring) (#62823 ) * add alerting prefix to some packages from alerting that have similar names in prometheus alertmanager	2023-02-03 11:36:49 -05:00
idafurjes	982939111b	Rename Id to ID for annotation models (#62886 ) * Rename Id to ID for annotation models * Add xorm tags * Rename Id to ID for API key models * Add xorm tags	2023-02-03 17:23:09 +01:00
Alexander Weaver	9eeea8f5ea	Alerting: Add label query parameters to state history endpoint (#62831 ) * Allow equality-only matching of arbitrary labels via query params * Pre-initialize map	2023-02-02 16:52:08 -06:00
Jean-Philippe Quéméner	1694a35e0c	Alerting: implement loki query for alert state history (#61992 ) * Alerting: implement loki query for alert state history * extract selector building * add unit tests for selector creation * backup * give selectors their own type * build dataframe * add some tests * small changes after manual testing * use struct client * golint * more golint * Make RuleUID optional for Loki implementation * Drop initial assumption that we only have one series * Pare down to three columns, fix timestamp overflows, improve failure cases in loki responses * Embed structred log lines in the dataframe as objects rather than json strings * Include state history label filter * Remove dead code --------- Co-authored-by: Alex Weaver <weaver.alex.d@gmail.com>	2023-02-02 16:31:51 -06:00
Matthew Jacobson	f9ec16e74f	Alerting: Fix template validation in provisioning api (#62530 ) * Alerting: Fix template validation in provisioning api Fix issue where provisioning API accepts a malformed template having extra text outside of definition block and template name matching definition name.	2023-02-02 15:26:39 -05:00
Alexander Weaver	647f73ddc5	Alerting: Add static label to all state history entries (#62817 ) * Add static label to all state history entries * Separate label and value visually	2023-02-02 13:25:26 -06:00
Santiago	ba731f7865	Alerting: Mark AM configuration as applied (#61330 ) * Mark AM configuration as applied * add missing checks, make linter happy * fix deadlock, mark as valid on save and on load * mark configurations only if needed * check error after applyConfig() * code review comments * code review changes * more code review changes * clean HistoricConfigFromAlertConfig function	2023-02-02 14:45:17 -03:00
Alexander Weaver	6ad1cfef38	Alerting: Add endpoint for querying state history (#62166 ) * Define endpoint and generate * Wire up and register endpoint * Cleanup, define authorization * Forgot the leading slash * Wire up query and SignedInUser * Wire up timerange query params * Add todo for label queries * Drop comment * Update path to rules subtree	2023-02-02 11:34:00 -06:00
Alexander Weaver	9fa28c11c5	Alerting: Usability adjustments to Loki representation of state history values (#62643 ) * Extract label merge, add test file * Extract error/NoData to first class fields, remove a layer from values * Include dashUID and panelID as line-level fields * Drop unnecessary object receiver * Add tests for stream building * Drop NoData field from log lines	2023-02-02 10:54:10 -06:00
idafurjes	23c27cffb3	Chore: Rename Id to ID in alerting models (#62777 ) * Chore: Rename Id to ID in alerting models * Add xorm tags for datasource * Add xorm tag for uid	2023-02-02 17:22:43 +01:00
Sonia Aguilar	753c84f825	Alerting: Pass yaml as a query param in export request (#62751 ) * Set YAML as default value for exporting alert rules * use YAML format for rule list export Co-authored-by: Sonia Aguilar <33540275+soniaAguilarPeiron@users.noreply.github.com> * lint * Add new format query param to swagger+docs * Fix broken test --------- Co-authored-by: Gilles De Mey <gilles.de.mey@gmail.com> Co-authored-by: Matt Jacobson <matthew.jacobson@grafana.com>	2023-02-02 16:10:02 +00:00
Steve Simpson	c44e9f6b71	Alerting: Add metrics around notification delivery. (#62778 ) This change exposes more metrics from the embedded Alertmanager, which are valuable for troubleshooting Alertmanager operation particularly in HA setups. ``` grafana_alerting_notifications_total grafana_alerting_notifications_failed_total grafana_alerting_notification_requests_total grafana_alerting_notification_requests_failed_total grafana_alerting_notification_latency_seconds grafana_alerting_nflog_gc_duration_seconds grafana_alerting_nflog_snapshot_duration_seconds grafana_alerting_nflog_snapshot_size_bytes grafana_alerting_nflog_queries_total grafana_alerting_nflog_query_errors_total grafana_alerting_nflog_query_duration_seconds grafana_alerting_nflog_gossip_messages_propagated_total grafana_alerting_dispatcher_aggregation_groups grafana_alerting_dispatcher_alert_processing_duration_seconds ``` Note that `alertmanager_dispatcher_aggregation_group_limit_reached_total` is explicitly not exposed, as the group limit metrics are not enabled.	2023-02-02 14:44:20 +01:00
Alexander Weaver	fcecf4d3cb	Alerting: Refactor away a layer of indirection around the goroutine in Loki state history (#62644 ) Inline recordStreamsAsync in loki backend	2023-02-01 11:57:29 -06:00
Gilles De Mey	26866953c1	Alerting: hide "silence" button for external AM setups (#62133 )	2023-02-01 15:51:05 +01:00
Sofia Papagiannaki	f143b0a5b2	Chore: Move folder store interface, implementation and test under pkg/services/folder (#62586 ) * Chore: Move folder store into folder service package * Split folder and dashboard store implementations	2023-02-01 15:43:21 +02:00
Alex Moreno	53945afedf	Alerting: Allow alert rule pausing from API (#62326 ) * Add is_paused attr to the POST alert rule group endpoint * Add is_paused to alerting API POST alert rule group * Fixed tests * Add is_paused to alerting gettable endpoints * Fix integration tests * Alerting: allow to pause existing rules (#62401) * Display Pause Rule switch in Editing Rule form * add isPaused property to form interface and dto * map isPaused prop with is_paused value from DTO Also update test snapshots * Append '(Paused)' text on alert list state column when appropriate * Change Switch styles according to discussion with UX Also adding a tooltip with info what this means * Adjust styles * Fix alignment and isPaused type definition Co-authored-by: gillesdemey <gilles.de.mey@gmail.com> * Fix test * Fix test * Fix RuleList test --------- Co-authored-by: gillesdemey <gilles.de.mey@gmail.com> * wip * Fix tests and add comments to clarify AlertRuleWithOptionals * Fix one more test * Fix tests * Fix typo in comment * Fix alert rule(s) cannot be paused via API * Add integration tests for alerting api pausing flow * Remove duplicated integration test --------- Co-authored-by: Virginia Cepeda <virginia.cepeda@grafana.com> Co-authored-by: gillesdemey <gilles.de.mey@gmail.com> Co-authored-by: George Robinson <george.robinson@grafana.com>	2023-02-01 13:15:03 +01:00
Alexander Weaver	03f3fbec0d	Alerting: Fix handling of special floating-point cases when writing observed values to annotations (#61074 ) * Fix json serialization of state values * Simplify two of the tests * Fix linter complaint * Don't return error if we fail to look up dashboard, just log it and move on * Address linter complaint	2023-01-31 15:27:51 -06:00
gotjosh	55e7cf1aed	Alerting: Introduce Metric Aggregation starting with Silences (#62512 ) * Alerting: Introduce Metric Aggregation starting with Silences --------- Co-authored-by: Alexander Weaver <weaver.alex.d@gmail.com>	2023-01-31 19:54:38 +00:00
gotjosh	178f290f0c	Update dskit to the latest main (#62616 ) * Update dskit to the latest main * Break free from a cortex depedency	2023-01-31 19:05:49 +00:00
ismail simsek	91221bc436	Expressions: Fixes the issue showing expressions editor (#62510 ) * Use suggested value for uid * update the snapshot * use __expr__ * replace all -100 with __expr__ * update snapshot * more changes * revert redundant change * Use expr.DatasourceUID where it's possible * generate files	2023-01-31 18:50:10 +01:00
Alexander Weaver	e7ace4ed62	Alerting: Allow separate read and write path URLs for Loki state history (#62268 ) Extract config parsing and add tests	2023-01-30 16:30:05 -06:00
Alexander Weaver	b4682fe3cb	Alerting: Configurable externalLabels for Loki state history (#62404 ) * Add config option for external labels * Remove redundant nilcheck	2023-01-30 14:24:45 -06:00
Alex Moreno	7a465f42a6	Alerting: Allow pausing alerts from provisioning (#62263 ) * Allow pausing alerts from provisioning * Update swagger * Add IsPaused to provision export endpoints * Add pause field in sample.yml * Add exception for reset state in first loop iteration of scheduler if rule is paused * Update provision definition and swagger docs * Fix provisioning export tests * Suggestion: Simplify if condition * Add more context to a comment	2023-01-30 16:29:05 +01:00
Ieva	ee3d742c7d	RBAC: inherit folder permissions when resolving managed permissions (#62244 ) * add nested folder scope inheritance to managed permission services * add a more specific erorr * remove circular dependencies * use errutil for returning erorr * fix tests * fix tests * define a new error in ac package	2023-01-30 14:19:42 +00:00
idafurjes	3bda112c5f	Chore: Move search model from models package to search service (#62215 ) * Chore: Move search model from models package to search service * Remove unused imports * Cleanup after merge	2023-01-30 15:17:53 +01:00
Serge Zaitsev	d6d4097567	Chore: Fix goimports grouping in alerting (#62424 ) * fix goimports * fix goimports order	2023-01-30 09:55:35 +01:00
Yuri Tseretyan	0c4671e31f	Alerting: Update historian to ignore transitions from Normal Paused and Updated (#62267 )	2023-01-27 16:26:22 -05:00
gotjosh	3c616da83f	Alerting: Refactor metrics/ngalert.go into seperate files (#62362 ) * Alerting: Refactor metrics/ngalert.go into seperate files	2023-01-27 18:49:49 +00:00
Matthew Jacobson	c006df375a	Alerting: Create endpoints for exporting in provisioning file format (#58623 ) This adds provisioning endpoints for downloading alert rules and alert rule groups in a format that is compatible with file provisioning. Each endpoint supports both json and yaml response types via Accept header as well as a query parameter download=true/false that will set Content-Disposition to recommend initiating a download or inline display. This also makes some package changes to keep structs with potential to drift closer together. Eventually, other alerting file structs should also move into this new file package, but the rest require some refactoring that is out of scope for this PR.	2023-01-27 11:39:16 -05:00
George Robinson	b6db1ed524	Revert "Alerting: Add is_paused attr to the POST alert rule group endpoint" (#62310 ) Revert "Alerting: Add is_paused attr to the POST alert rule group endpoint (#62253)" This reverts commit `3ccafe3a5a`.	2023-01-27 13:41:36 +01:00
Alex Moreno	3ccafe3a5a	Alerting: Add is_paused attr to the POST alert rule group endpoint (#62253 ) Add is_paused attr to the POST alert rule group endpoint	2023-01-27 10:50:06 +01:00
Yuri Tseretyan	05bf241952	Alerting: Update state manager to return StateTransitions when Delete or Reset (#62264 ) * update Delete and Reset methods to return state transitions this will be used by notifier code to decide whether alert needs to be sent or not. * update scheduler to provide reason to delete states and use transitions * update FromAlertsStateToStoppedAlert to accept StateTransition and filter by old state * fixup * fix tests	2023-01-27 09:46:21 +01:00
idafurjes	6c5a573772	Chore: Move ReqContext to contexthandler service (#62102 ) * Chore: Move ReqContext to contexthandler service * Rename package to contextmodel * Generate ngalert files * Remove unused imports	2023-01-27 08:50:36 +01:00
Alex Moreno	531b439cf1	Alerting: Add alert pausing feature (#60734 ) * Add field in alert_rule model, add state to alert_instance model, and state to eval * Remove paused state from eval package * Skip paused alert rules in scheduler * Add migration to add is_paused field to alert_rule table * Convert to postable alerts only if not normal, pernding, or paused * Handle paused eval results in state manager * Add Paused state to eval package * Add paused alerts logic in scheduler * Skip alert on scheduler * Remove paused status from eval package * Apply suggestions from code review Co-authored-by: George Robinson <george.robinson@grafana.com> * Remove state * Rethink schedule and manager for paused alerts * Change return to continue * Remove unused var * Rethink alert pausing * Paused alerts storing annotations * Only add one state transition * Revert boolean method renaming refactor * Revert take image refactor * Make registry errors public * Revert method extraction for getting a folder title * Revert variable renaming refactor * Undo unnecessary changes * Revert changes in test * Remove IsPause check in PatchPartiLAlertRule function * Use SetNormal to set state * Fix text by returning to old behaviour on alert rule deletion * Add test in schedule_unit_test.go to test ticks with paused alerts * Add coment to clarify usage of context.Background() * Add comment to clarify resetStateByRuleUID method usage * Move rule get to a more limited scope * Update pkg/services/ngalert/schedule/schedule.go Co-authored-by: George Robinson <george.robinson@grafana.com> * rum gofmt on pkg/services/ngalert/schedule/schedule.go * Remove defer cancel for context * Update pkg/services/ngalert/models/instance_test.go Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> * Update pkg/services/ngalert/models/testing.go Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> * Update pkg/services/ngalert/schedule/schedule_unit_test.go Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> * Update pkg/services/ngalert/schedule/schedule_unit_test.go Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> * Update pkg/services/ngalert/models/instance_test.go Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> * skip scheduler rule state clean up on paused alert rule * Update pkg/services/ngalert/schedule/schedule.go Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> * Fix mock in test * Add (hopefully) final suggestions * Use error channel from recordAnnotationsSync to cancel context * Run make gen-cue * Place pause alert check in channel update after version check * Reduce branching un update channel select * Add if for error and move code inside if in state manager ResetStateByRuleUID * Add reason to logs * Update pkg/services/ngalert/schedule/schedule.go Co-authored-by: George Robinson <george.robinson@grafana.com> * Do not delete alert rule routine, just exit on eval if is paused * Reduce branching and create-close a channel to avoid deadlocks * Separate state deletion and state reset (includes history saving) * Add current pause state in rule route in scheduler * Split clearState and bring errCh closer to RecordStatesAsync call * Change rule to ruleMeta in RecordStatesAsync * copy state to be able to modify it * Add timeout to context creation * Shorten the timeout * Use resetState is rule is paused and deleteState if rule is not paused * Remove Empty state reason * Save every rule change in historian * Add tests for DeleteStateByRuleUID and ResetStateByRuleUID * Remove useless line * Remove outdated comment Co-authored-by: George Robinson <george.robinson@grafana.com> Co-authored-by: Santiago <santiagohernandez.1997@gmail.com> Co-authored-by: Armand Grillet <2117580+armandgrillet@users.noreply.github.com>	2023-01-26 18:29:10 +01:00
Kristin Laemmert	e8b8a9e276	chore: move dashboard_acl models into dashboard service (#62151 )	2023-01-26 08:46:30 -05:00
gotjosh	0bfe150928	Alerting: Fix Test Receivers when settings are non-strings (#62156 ) * Alerting: Fix Test Receivers when settings are non-strings As part of the Alerting extraction, we want to make sure we don't have circular depedencies. As such, I had to move `PostableGrafanaReceiver` to a new struct in `grafana/alerting` called `GrafanaReceiver`. `PostableGrafanaReceiver` has an attribute called `Settings` that uses a Grafana-propietary struct called `RawMessage`, this struct shadows `json.RawMessage`. When I created `GrafanaReceiver`, I turned settings into a `map[string]string` thinking all settings would end up as strings. This was a mistake, and this test proves that it doesn't work, and breaks the API.	2023-01-26 12:54:03 +00:00
George Robinson	a7eab8e46e	Alerting: Support context.Context in Loki interface (#61979 ) This commit adds support for canceleable contexts in the Loki interface.	2023-01-26 09:31:20 +00:00
Sofia Papagiannaki	cd27562c76	Access control: Modify dashboard/folder resolvers so that return also the inherited scopes (#62025 ) * Access Control: Add folder service dependency to the dashboard/folder resolvers * Expose the function fetching parents to folder interface * Add generic prepend utility * Modify dashboard resolvers to return inherited scopes	2023-01-26 10:21:10 +02:00
Alexander Weaver	eb1293ebb1	Alerting: Re-generate swagger definitions (#62154 ) Regenerate swagger, add body binding so parameters work	2023-01-25 15:01:42 -06:00
Alexander Weaver	046a9bb7c1	Alerting: Copy rule definitions into state history (#62032 ) * Copy rules instead of accepting pointer * Deep-copy the rule, for even more guarantees * Create struct just for needed fields * Move RuleMeta to historian/model package, iron out package dependencies * Move tests for dash ID parsing to model package along with code	2023-01-25 11:29:57 -06:00
idafurjes	b54b80f473	Chore: Remove Result from dashboard models (#61997 ) * Chore: Remove Result from dashboard models * Fix lint tests * Fix dashboard service tests * Fix API tests * Remove commented out code * Chore: Merge main - cleanup	2023-01-25 10:36:26 +01:00
idafurjes	421976e919	Chore: Remove folders from models pkg (#61853 )	2023-01-25 09:14:32 +01:00
Santiago	e5920c211e	Chore: Fix random indices for slices in test files (#61884 ) * Fix random indices for slices in test files * Empty commit	2023-01-24 15:07:37 -03:00
George Robinson	239d94205a	Alerting: Return chan <-error for #61811 (#61858 )	2023-01-24 15:41:38 +00:00
Alexander Weaver	7ccc845187	Alerting: Push state history entries to Loki (#61724 ) * Implement push endpoint * Drop duplicated struct * Genericize auth/tenant headers and improve logging in error case * Flesh out the data model * Drop dead code * Drop log line entirely * Drop unused arg * Rename a few type manipulation functions * Extract label keys as constants * Improve logs when loki responds with error * Inline lokiRepresentation function	2023-01-23 16:31:03 -06:00
Sofia Papagiannaki	c7a7ebd3e0	Chore: Drop search service dependency from folder service (#61789 ) * Chore: Drop search service dependency from folder service	2023-01-23 14:09:09 +02:00
gotjosh	0be920e61c	Alerting: Remove unused code after importing from grafana/alerting (#61869 ) * Alerting: Remove unused code after importing from grafana/alerting	2023-01-23 10:30:10 +00:00
gotjosh	511dab3b4b	Update `grafana/alerting` to the latest main (#61810 ) * Update `grafana/alerting` to the latest main Also updates prometheus-alertmanager since we use that one directly for some structs.	2023-01-19 20:44:49 +00:00
Sofia Papagiannaki	c104cc7020	Chore: Split folder store and dashboard store interfaces (#61655 ) * update folder store mock * Split folder store and dashboard store interfaces	2023-01-19 18:38:07 +02:00
Alexander Weaver	c10713ea76	Alerting: Create query interface for state history along with annotation-based implementation (#61646 )	2023-01-19 10:45:31 +01:00
Yuri Tseretyan	2c46f46d37	Alerting: Rule evaluator to get cached data source info (#61305 ) do not skip cache when get data source info	2023-01-18 14:25:11 -05:00
Jean-Philippe Quéméner	44b11d3228	Alerting: support basic auth for the state history loki client (#61696 )	2023-01-18 20:24:40 +01:00
George Robinson	d4256b352d	Docs: Rename Message templates to Notification templates (#59477 ) This commit renames "Message templates" to "Notification templates" in the user interface as it suggests that these templates cannot be used to template anything other than the message. However, message templates are much more general and can be used to template other fields too such as the subject of an email, or the title of a Slack message.	2023-01-18 17:26:34 +00:00
Sofia Papagiannaki	b80c9bb974	Chore: Drop dashboard service dependency from folder service (#61614 ) * Chore: Drop dashboard dependency from folder service	2023-01-18 17:47:59 +02:00
Santiago	b5fa9e3501	Chore: Fix "manger" typo (#61649 ) fix mangers -> managers	2023-01-17 23:13:27 +00:00
Alexander Weaver	1ac89ea040	Alerting: Add client configuration for remote Loki historian backend and test connection (#61114 ) * Create loki client type and ping method * Expose TestConnection on client * Configure and ping Loki URL * Close response body reader if present * Add 30 second timeout * Remove duplicate close	2023-01-17 13:58:52 -06:00
Kristin Laemmert	f6e3252c00	chore: move notifications models into notifications service (#61638 )	2023-01-17 14:47:31 -05:00
Matthew Jacobson	23e05373a7	Alerting: Fix flaky TestIntegrationUpdateAlertRules (#61641 ) Prevents random OrgID=0 in test alert generation causing invalid alert rule.	2023-01-17 19:09:46 +00:00
Alexander Weaver	4f1bdc0607	Alerting: Skip flaky test in TestIntegrationUpdateAlertRules (#61627 ) * Skip flaky test * Add comment	2023-01-17 10:39:16 -06:00
Denis Limarev	e6dee8a723	Perfomance: Preallocate slices (#61580 )	2023-01-17 11:50:17 +00:00
idafurjes	7c2522c477	Chore: Move dashboard models to dashboard pkg (#61458 ) * Copy dashboard models to dashboard pkg * Use some models from current pkg instead of models * Adjust api pkg * Adjust pkg services * Fix lint	2023-01-16 16:33:55 +01:00
Jo	dcfeab2c73	AuthN: User Quota (#61540 ) * remove reqContext from quota checks in login * add guards for nil ScopeParams	2023-01-16 11:54:15 +01:00
Yuri Tseretyan	9d57b1c72e	Alerting: Do not persist noop transition from Normal state. (#61201 ) * add feature flag `alertingNoNormalState` * update instance database to support exclusion of state in list operation * do not save normal state and delete transitions to normal * update get methods to filter out normal state	2023-01-13 18:29:29 -05:00
Ghazanfar	d553a016cc	Alerting: UI changes required to support v3 and Auth in Kafka Contact Point (#61123 )	2023-01-13 17:45:43 -05:00
Alexander Weaver	b289b8ac6e	Alerting: Set error annotation on EvaluationError regardless of underlying error type (#61506 ) Set error annotation regardless of underlying error type	2023-01-13 13:58:02 -06:00
gotjosh	e7cd6eb13c	Alerting: Use `alerting.GrafanaAlertmanager` instead of initialising Alertmanager components directly (#61230 ) * Alerting: Use `alerting.GrafanaAlertmanager` instead of initialising Alertmanager components directly	2023-01-13 12:54:38 -04:00
gotjosh	8f72893076	Alerting: Document not supporting inhibition rules (#61313 ) * Alerting: Document not supporting inhibition rules * Update docs/sources/alerting/manage-notifications/create-silence.md Co-authored-by: brendamuir <100768211+brendamuir@users.noreply.github.com> * Update docs/sources/alerting/manage-notifications/alertmanager.md Co-authored-by: brendamuir <100768211+brendamuir@users.noreply.github.com> Co-authored-by: brendamuir <100768211+brendamuir@users.noreply.github.com>	2023-01-13 15:06:06 +01:00
gotjosh	49ae1bbe63	Introduce `AlertingConfiguration` that implements `alerting.Configuration` (#61427 ) * Introduce `AlertingConfiguration` that implements `alerting.Configuration`	2023-01-12 16:03:07 -04:00
gotjosh	fd6f107ded	Alerting Unification: Use the errors from grafana/alerting in Alerts (#61425 )	2023-01-12 15:23:34 -04:00
gotjosh	ddb85ad6ad	Use the `ClusterPeer` interface from grafana/alerting (#61409 ) * Use the Cluster interface from grafana/alerting	2023-01-12 14:47:22 -04:00
gotjosh	2d1faae0b5	Alerting Unification: Use `alerting.MaintenanceOptions` to configure silences and nflog (#61384 )	2023-01-12 12:31:38 -04:00
gotjosh	39e429a14b	Alerting Unification: Use the errors from grafana/alerting in Silences (#61334 )	2023-01-12 10:03:49 -04:00
gotjosh	f85a948214	Alerting Unification: Use the State interface from the alerting package (#61333 )	2023-01-11 19:50:45 -04:00
Denis Limarev	90badc8729	Performance: Add preallocation for some slices (#59593 )	2023-01-11 18:03:37 +01:00
Yuri Tseretyan	b4e1e1871f	Alerting: Fix evaluation timeout (#61303 )	2023-01-11 10:52:54 -05:00
Yuri Tseretyan	86b5fbbf60	Alerting: Introduce state manager config structure (#61249 )	2023-01-10 16:26:15 -05:00
George Robinson	2a291afbae	Alerting: Use consts from alerting package (#61241 )	2023-01-10 19:59:13 +00:00
George Robinson	d19d8c6625	Alerting: Update Alerting and Alertmanager to v0.25.1 (#61233 ) Update Alerting and Alertmanager to v0.25.1	2023-01-10 16:17:07 +00:00
Yuri Tseretyan	da18c89e91	Alerting: Scheduler to call DeleteAlertRule once when it stops deleted rules (#61189 ) scheduler to call DeleteAlertRule once when it stops deleted rules	2023-01-09 14:39:32 -05:00
Yuri Tseretyan	48f1db63ff	Alerting: Add support for tracing to alerting scheduler (#61057 )	2023-01-06 21:21:43 -05:00
Alexander Weaver	eb960d9725	Alerting: Add un-documented toggle for changing state history backend, add shells for remote loki and sql (#61072 ) * Add toggle for state history backend and shells * Extract some shared logic and add tests	2023-01-06 12:06:01 -06:00
Alexander Weaver	8c3a5f6da0	Alerting: Allow state history to be disabled through configuration (#61006 ) * Add configuration option for if state history should be enabled * Inject no-op when history is disabled	2023-01-05 12:21:07 -06:00
George Robinson	9af7adef76	Alerting: Support customizable timeout for screenshots (#60981 ) This commit adds a customizable timeout for screenshots called capture_timeout. The default value is 10 seconds, and the maximum value is 30 seconds. This timeout should be less than the minimum Interval of all Evaluation Groups to avoid back pressure on alert rule evaluation.	2023-01-05 16:07:46 +00:00
Alexander Weaver	0e7640475f	Alerting: Store alertmanager configuration history in a separate table in the database (#60492 ) * Update config store to split between active and history tables * Migrations to fix up indexes * Implement migration from old format to new * Move add migrations call * Delete duplicated rows * Explicitly map fields * Quote the column name because it's a reserved word * Lift migrations to top * Use XORM for nearly everything, avoid any non trivial raw SQL * Touch up indexes and zero out IDs on move * Drop TODO that's already completed * Fix assignment of IDs	2023-01-04 10:43:26 -06:00
Yuri Tseretyan	4d989860fb	Alerting: Fix conversion of alert state from db state during manager warmup (#60933 )	2023-01-04 09:40:04 -05:00
Alexander Weaver	b88b8bc291	Alerting: Fix missing dashboard/panelID links in annotations (#60926 ) Assign thru ref	2023-01-03 14:12:27 -06:00
Santiago	05c9af5110	Extract custom template functions (#60695 ) extract custom template functions and export the FuncMap	2022-12-22 17:31:40 -03:00
Yuri Tseretyan	f990be58cb	Alerting: Use all notifiers from alerting repository (#60655 )	2022-12-22 09:27:18 -05:00
Marcus Efraimsson	c35c689a96	Plugins: Automatically forward plugin request HTTP headers in outgoing HTTP requests (#60417 ) Automatically forward core plugin request HTTP headers in outgoing HTTP requests. Core datasource plugin authors don't have to specifically handle forwarding of HTTP headers, e.g. do not have to "hardcode" the header-names in the datasource plugin, if not having custom needs. Fixes #57065	2022-12-21 13:25:58 +01:00
Yuri Tseretyan	dc2ca80f4d	Alerting: Refactor email notifier (#60602 ) * refactor email to not use simplejson * add tests * split integration test and unit test + more unit-tests * Remove outdated comment Co-authored-by: Armand Grillet <2117580+armandgrillet@users.noreply.github.com>	2022-12-21 02:03:15 -05:00
Yuri Tseretyan	4a3097f52a	Alerting: Update Discord receiver to use encoding/json to build a webhook message + truncate long message (#60592 ) * replace simplejson with models * truncate too long messages Co-authored-by: Santiago <santiagohernandez.1997@gmail.com>	2022-12-20 14:20:42 -05:00
Yuri Tseretyan	aaa55b4252	Alerting: Update Kafka receiver to use encoding/json to build messages (#60593 ) Co-authored-by: Santiago <santiagohernandez.1997@gmail.com>	2022-12-20 14:20:09 -05:00
Yuri Tseretyan	a0bf62cc9e	Alerting: Update receivers to use app version from factory config (#60585 )	2022-12-20 11:23:10 -05:00
Yuri Tseretyan	ec45c9c990	Alerting: update dingding, discord, googlechat, kafka, line notifiers to use encoding/json to parse settings (#60542 ) also, rename Content to Message to match JSON name for Discord and GoogleChat	2022-12-20 09:46:13 -05:00
Yuri Tseretyan	35090c376c	Alerting: Replace VictorOps receiver with the one from alerting repository (#60543 ) * replace victorops with one from alerting * update other usages	2022-12-20 10:55:41 +01:00
Alexander Weaver	ca3f8ba6f4	Alerting: Refactor alertmanager notifier to use encoding/json to parse settings instead of simplejson (#55507 ) * replace basic auth header with method call Co-authored-by: Yuri Tseretyan <yuriy.tseretyan@grafana.com>	2022-12-19 15:12:49 -05:00
Yuri Tseretyan	f0cabe14d5	Alerting: import Grafana alerting package and update usages (#60490 ) * update remaining notifiers to use alerting package	2022-12-19 10:53:58 -05:00
Yuri Tseretyan	92d12fdefa	Alerting: Remove fake secret service in tests (#60488 )	2022-12-16 15:01:41 -05:00
Yuri Tseretyan	9ad45aedcf	Alerting: replace usage of simplejson to json.RawMessage in NotificationChannelConfig (#60423 ) * introduce alias for json.RawMessage with name RawMessage. This is needed to keep raw JSON and implement a marshaler for YAML, which does not seem to be used but there are tests that fail. * replace usage of simplejson with RawMessage in NotificationChannelConfig * remove usage of simplejson in tests * change migration code to convert simplejson to raw message	2022-12-16 13:01:06 -05:00
Alexander Weaver	91bd1cdb41	Revert "Alerting: Store alertmanager configuration history in a separate table in the database" (#60470 ) Revert "Alerting: Store alertmanager configuration history in a separate table in the database (#60197)" This reverts commit `ec80f38c34`.	2022-12-16 10:07:44 -05:00
Alex Moreno	174c61b949	Alerting: Set Dashboard and Panel IDs on rule group replacement (#60374 ) * Set Dashboard and Panel IDs on rule group replacement * fix comments and abbreviate test variable name * Update pkg/services/ngalert/provisioning/alert_rules.go Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com> Co-authored-by: Jean-Philippe Quéméner <JohnnyQQQQ@users.noreply.github.com>	2022-12-16 11:47:25 +01:00
Alexander Weaver	ec80f38c34	Alerting: Store alertmanager configuration history in a separate table in the database (#60197 ) * Update config store to split between active and history tables * Migrations to fix up indexes * Implement migration from old format to new * Move add migrations call * Delete duplicated rows * Explicitly map fields * Quote the column name because it's a reserved word * Lift migrations to top	2022-12-15 17:35:00 -06:00
Yuri Tseretyan	6637333748	Alerting: refactor notifiers to use package specific Logger interface (#60361 ) * introduce Logger interface local to channles + implementaton that wraps the Grafana logger * make NewFactoryConfig accept LoggerFactory * add logger field to FactoryConfig * update usages of log.Logger to internal interface	2022-12-15 11:10:31 -05:00
Sofia Papagiannaki	11d8bcbea9	Guardian: Introduce additional constructors (#59577 ) * Guardian: Use dashboard UID instead of ID * Apply suggestions from code review Introduce several guardian constructors and each time use the most appropriate one.	2022-12-15 16:34:17 +02:00
Yuri Tseretyan	0e7c95a4d2	Alerting: Remove reference to global models package in channels package (#60358 ) * remove intermediate struct to create Base struct * fix alertmanager	2022-12-14 16:21:55 -05:00
Kristina	5a7f38053b	Remove explore compact URLs (#59686 ) * Remove explore compact URLs * Remove two explore link builders that create compact URLs * Fix merge conflict	2022-12-14 12:57:53 -06:00
Yuri Tseretyan	de008005ce	Alerting: isolate ImageStore in notify package (#60353 )	2022-12-14 13:20:20 -05:00
Yuri Tseretyan	7c3ab4a715	Alerting: Remove dependency on Grafana notifications package in alerting notifiers (#60271 ) * create sender service interface and bridge to grafana notifier service * update notifiers to use local sender interface	2022-12-14 10:59:37 -05:00
Yuri Tseretyan	07b5043222	Alerting: Add support for settings parse_mode and disable_notifications to Telegram reciever (#60198 )	2022-12-14 10:44:39 -05:00
Yuri Tseretyan	ad09feed83	Alerting: rule backtesting API (#57318 ) * Implement backtesting engine that can process regular rule specification (with queries to datasource) as well as special kind of rules that have data frame instead of query. * declare a new API endpoint and model * add feature toggle `alertingBacktesting`	2022-12-14 09:44:14 -05:00
Alexander Weaver	821614fb43	Alerting: Align notifier truncation and logging with prometheus/alertmanager (#59339 ) * Move truncation code to util to mirror upstream * Resolve merge conflicts * Align logging of alert key * Update tests and fix field passing bug * Remove superfluous newline in test now that we trim whitespace * Uptake minor log changes from upstream	2022-12-13 19:50:24 -06:00
Alexander Weaver	e97b43cd58	Alerting: Add provisioning endpoint to fetch all rules (#59989 ) * Domain layer api for fetching all rules * Add endpoint for getting all rules	2022-12-13 11:54:08 +01:00
Alexander Weaver	595e623c28	Alerting: Additional tests for the config store (#60130 ) Additional tests for the config store	2022-12-12 11:11:18 -06:00
Yuri Tseretyan	df7f636759	Alerting: Fix slack receiver to close file descriptors when they're not needed anymore (#60178 )	2022-12-12 11:19:02 -05:00
Yuri Tseretyan	4374966987	Alerting: Replace hardcoded <no value> to [no value] in label expansion (#60129 ) * replace hardcoded <no value> to [no value] in label expansion	2022-12-12 10:12:30 -05:00
Joe Blubaugh	1a8d0e2736	Alerting: Speed up unit and integration tests. (#60067 ) This change marks tests in the `sender` package that use an external process as integration tests instead of unit tests. This speeds up the package's unit tests by about 20 seconds. This change also reduces the number of alert instances in the `store` package's bulk write integration test from 20_000 to 10_000. This is still enough to exercise the bulk-write code but speeds up the package tests from about 250s to 130s. Put together, integration tests go to about 160s while also speeding up unit tests by 20s.	2022-12-12 14:21:06 +08:00
George Robinson	76601f3ae7	Alerting: Better define how we set states (#59977 ) This commit better defines how we set states in resultNormal, resultAlerting, resultError and resultNoData. It changes the existing code to call methods such as SetAlerting, SetPending, SetNormal, SetError and NoData instead of assigning values to each individual field whenever the state is changed. This should make it easier to understand what fields should be set for which states and avoid cases where states are missing, or have additional unexpected fields.	2022-12-08 20:12:13 +00:00
Yuri Tseretyan	316870c658	Alerting: PagerDuty receiver to let user configure fields Source, Client and Client URL (#59895 ) * add support for source field * add client_url * use real host name for source placeholder	2022-12-08 11:49:27 -05:00
Joe Blubaugh	e6743a7e9a	Alerting: Use the QuotaTargetSrv instead of the QuotaTarget in quota check (#60026 ) Before this change, the alerting provisioning system incorrectly used the QuotaTarget to check if alerting's request quota had been reached. The quota service requires the QuotaTargetSrv, which is what's registered with the service at startup time. This is leading to errors in the provisioning system.	2022-12-08 22:34:46 +08:00
Yuri Tseretyan	c5ee4e4ae1	Alerting: Improve rule validation to check if rule uses backend datasources (#58986 ) * validate if rule uses backend datasources * add backend datasource to test * fix tests * another forgotten import * remove unused var	2022-12-08 10:44:02 +01:00
George Robinson	6359dab040	Alerting: Change resultError in preparation for supporting ForError duration (#59894 )	2022-12-07 10:45:56 +00:00
Serge Zaitsev	43f40e6c7c	Chore: Replace yaml.v2 with yaml.v3 (#59897 ) * replace yaml.v2 with yaml.v3 * fix a few tests due to the yaml.v3 api changes * and another goconvey mistake in tests	2022-12-06 21:17:17 +01:00
George Robinson	3c249e1b99	Fix incorrect start time for DatasourceError alerts (#59903 )	2022-12-06 18:44:06 +00:00
Yuri Tseretyan	abb49d96b5	Alerting: update state manager to return StateTransition instead of State (#58867 ) * improve test for stale states * update state manager return StateTransition * update scheduler to accept state transitions	2022-12-06 13:07:39 -05:00

... 2 3 4 5 6 ...

1151 Commits