elasticsearch

Commit Graph

Author	SHA1	Message	Date
David Roberts	cbb8b17d74	[DOCS] Docs changes for overridden delimiter in find_file_structure (#56288 ) Docs for #55735 Co-authored-by: Lisa Cawley <lcawley@elastic.co>	2020-05-14 09:24:07 +01:00
Lisa Cawley	84e28e42c8	[DOCS] Clarify model snapshot retention properties (#56477 )	2020-05-11 07:41:47 -07:00
István Zoltán Szabó	c994369893	[DOCS] Expands GET DFA stats API docs with new phases (#56407 ) Co-authored-by: Lisa Cawley <lcawley@elastic.co>	2020-05-11 09:22:30 +02:00
David Roberts	c99021cdcb	[ML] More advanced model snapshot retention options (#56125 ) This PR implements the following changes to make ML model snapshot retention more flexible in advance of adding a UI for the feature in an upcoming release. - The default for `model_snapshot_retention_days` for new jobs is now 10 instead of 1 - There is a new job setting, `daily_model_snapshot_retention_after_days`, that defaults to 1 for new jobs and `model_snapshot_retention_days` for pre-7.8 jobs - For days that are older than `model_snapshot_retention_days`, all model snapshots are deleted as before - For days that are in between `daily_model_snapshot_retention_after_days` and `model_snapshot_retention_days` all but the first model snapshot for that day are deleted - The `retain` setting of model snapshots is still respected to allow selected model snapshots to be retained indefinitely Closes #52150	2020-05-05 12:55:50 +01:00
Dimitris Athanasiou	6bf3834059	[ML] Add loss_function to regression (#56118 ) Adds parameters `loss_function` and `loss_function_parameter` to regression.	2020-05-05 12:36:05 +03:00
István Zoltán Szabó	86032ac56a	[DOCS] Simplifies footnote text in DFA APIs (#56105 ) Co-authored-by: Lisa Cawley <lcawley@elastic.co>	2020-05-05 09:03:16 +02:00
Lisa Cawley	52a2f7689f	[DOCS] Synchs and links hyperparameter descriptions (#55827 )	2020-05-04 07:37:14 -07:00
Lisa Cawley	5ef7aacbf7	[DOCS] Adds documentation for secondary authorization headers (#55365 ) Co-authored-by: Tim Vernum <tim@adjective.org>	2020-04-29 08:28:42 -07:00
István Zoltán Szabó	d70cef3474	[DOCS] Makes the footnotes less verbose in configuring aggs page. (#55857 )	2020-04-29 09:50:41 +02:00
István Zoltán Szabó	ca2f98382f	[DOCS] Changes feature importance links to point to the new page (#55531 ) * [DOCS] Changes feature importance links to point to the new page. * [DOCS] Fixes line breaks.	2020-04-28 09:02:14 +02:00
David Roberts	dcb6ed03cd	[ML] Adding failed_category_count to model_size_stats (#55716 ) The failed_category_count statistic records the number of times categorization wanted to create a new category but couldn't because the job had reached its model_memory_limit. Relates elastic/ml-cpp#1130	2020-04-25 08:01:21 +01:00
Lisa Cawley	7fafec0f8f	[DOCS] Update example and nesting in get data frame analytics job stats API (#55191 ) Co-Authored-By: Valeriy Khakhutskyy <1292899+valeriy42@users.noreply.github.com>	2020-04-22 08:07:31 -07:00
David Roberts	d1a9b3a545	[ML] Add effective max model memory limit to ML info (#55529 ) The ML info endpoint returns the max_model_memory_limit setting if one is configured. However, it is still possible to create a job that cannot run anywhere in the current cluster because no node in the cluster has enough memory to accommodate it. This change adds an extra piece of information, limits.effective_max_model_memory_limit, to the ML info response that returns the biggest model memory limit that could be run in the current cluster assuming no other jobs were running. The idea is that the ML UI will be able to warn users who try to create jobs with higher model memory limits that their jobs will not be able to start unless they add a bigger ML node to their cluster. Relates elastic/kibana#63942	2020-04-22 11:36:58 +01:00
David Roberts	8906e76079	[ML] Return assigned node in start/open job/datafeed response (#55473 ) Adds a "node" field to the response from the following endpoints: 1. Open anomaly detection job 2. Start datafeed 3. Start data frame analytics job If the job or datafeed is assigned to a node immediately then this field will return the ID of that node. In the case where a job or datafeed is opened or started lazily the node field will contain an empty string. Clients that want to test whether a job or datafeed was opened or started lazily can therefore check for this. Fixes #54067	2020-04-22 08:44:57 +01:00
István Zoltán Szabó	f8bfab2dab	[DOCS] Provides further details on aggregations in datafeeds (#55462 ) Co-authored-by: Lisa Cawley <lcawley@elastic.co>	2020-04-22 08:53:34 +02:00
Benjamin Trent	f72b2db29a	[ML] partitions model definitions into chunks (#55260 ) This paves the data layer way so that exceptionally large models are partitioned across multiple documents. This change means that nodes before 7.8.0 will not be able to use trained inference models created on nodes on or after 7.8.0. I chose the definition document limit to be 100. This SHOULD be plenty for any large model. One of the largest models that I have created so far had the following stats: ~314MB of inflated JSON, ~66MB when compressed, ~177MB of heap. With the chunking sizes of `16 * 1024 * 1024` its compressed string could be partitioned to 5 documents. Supporting models 20 times this size (compressed) seems adequate for now.	2020-04-20 15:13:23 -04:00
Benjamin Trent	c1afda4a23	[ML] adding prediction_field_type to inference config (#55128 ) Data frame analytics dynamically determines the classification field type. This field type then dictates the encoded JSON that is written to Elasticsearch. Inference needs to know about this field type so that it may provide the EXACT SAME predicted values as analytics. Here is added a new field `prediction_field_type` which indicates the desired type. Options are: `string` (DEFAULT), `number`, `boolean` (where close_to(1.0) == true, false otherwise). Analytics provides the default `prediction_field_type` when the model is created from the process.	2020-04-15 08:32:48 -04:00
Lisa Cawley	1f0341db39	[DOCS] Removes unshared sections from ml-shared.asciidoc (#55129 )	2020-04-14 15:19:31 -07:00
Lisa Cawley	998a085c14	[DOCS] Edits create data frame analytics job API (#54751 )	2020-04-13 09:58:03 -07:00
István Zoltán Szabó	b1b067c5ba	[DOCS] Adds link points to the data frame analytics supported fields (#55004 ) Co-authored-by: lcawl <lcawley@elastic.co>	2020-04-09 11:16:13 -07:00
István Zoltán Szabó	bb44726ad6	[DOCS] Reworks some parts of EMM API docs (#54872 ) Co-authored-by: Lisa Cawley <lcawley@elastic.co>	2020-04-08 09:50:12 +02:00
Lisa Cawley	c355fea8f4	[DOCS] Remove text fields from classification dependent variables (#54849 )	2020-04-07 10:43:15 -07:00
István Zoltán Szabó	1ae8bde756	[DOCS] Changes kibana_user to kibana_admin in DFA API prerequisites. (#54806 )	2020-04-06 15:45:08 +02:00
István Zoltán Szabó	a0662399c7	[DOCS] Makes PUT inference API docs collapsible (#54653 ) Co-authored-by: lcawl <lcawley@elastic.co>	2020-04-03 09:45:42 +02:00
Benjamin Trent	4e1ff31c3c	[ML] add new inference_config field to trained model config (#54421 ) A new field called `inference_config` is now added to the trained model config object. This new field allows for default inference settings from analytics or some external model builder. The inference processor can still override whatever is set as the default in the trained model config.	2020-04-02 10:34:17 -04:00
Benjamin Trent	1d24960ff8	[ML] prefer secondary authorization header for data[feed\|frame] authz (#54121 ) Secondary authorization headers are to be used to facilitate Kibana spaces support + ML jobs/datafeeds. Now on PUT/Update/Preview datafeed, and PUT data frame analytics the secondary authorization is preferred over the primary (if provided). closes https://github.com/elastic/elasticsearch/issues/53801	2020-04-02 10:10:46 -04:00
Benjamin Trent	bbd6e943de	[ML] add num_matches and preferred_to_categories to category defintion objects (#54214 ) This adds two new fields to category definitions. - `num_matches` indicating how many documents have been seen by this category - `preferred_to_categories` indicating which other categories this particular category supersedes when messages are categorized. These fields are only guaranteed to be up to date after a `_flush` or `_close` native change: https://github.com/elastic/ml-cpp/pull/1062	2020-04-02 07:49:09 -04:00
István Zoltán Szabó	b0f6d4ee0e	[DOCS] Updates estimate model memory docs (#54574 )	2020-04-01 15:53:53 +02:00
István Zoltán Szabó	b96743cfc5	[DOCS] Adds data_counts object to the GET DFA stats API (#54498 )	2020-04-01 10:05:00 +02:00
Jason Tedor	95a7eed9aa	Rename MetaData to Metadata in all of the places (#54519 ) This is a simple naming change PR, to fix the fact that "metadata" is a single English word, and for too long we have not followed general naming conventions for it. We are also not consistent about it, for example, METADATA instead of META_DATA if we were trying to be consistent with MetaData (although METADATA is correct when considered in the context of "metadata"). This was a simple find and replace across the code base, only taking a few minutes to fix this naming issue forever.	2020-03-31 15:52:01 -04:00
Lisa Cawley	b90e491f68	[DOCS] Collapses nested objects in data frame analytics APIs (#54472 )	2020-03-31 10:56:48 -07:00
Dimitris Athanasiou	5a98fc20e1	[ML] Fix DF analytics explain API request in docs (#54510 ) The explain API expects a data frame analytics config as its request.	2020-03-31 18:37:19 +03:00
István Zoltán Szabó	85d9b34dc5	[DOCS] Adds description of analysis_stats object and its properties to GET DFA stats API docs (#53881 ) Co-authored-by: Valeriy Khakhutskyy <1292899+valeriy42@users.noreply.github.com> Co-authored-by: Lisa Cawley <lcawley@elastic.co>	2020-03-31 13:27:54 +02:00
Lisa Cawley	fdcd19483d	[DOCS] Collapses content in machine learning APIs (#54234 )	2020-03-30 10:08:38 -07:00
Jason Tedor	1fc0432b24	Introduce formal role for remote cluster client (#53924 ) This commit introduce a formal role for identifying nodes that are capable of making connections to remote clusters.	2020-03-24 19:21:56 -04:00
David Roberts	8ee770560a	[ML] Add a model memory estimation endpoint for anomaly detection (#53507 ) A new endpoint for estimating anomaly detection job model memory requirements: POST _ml/anomaly_detectors/estimate_model_memory Closes #53219	2020-03-24 21:38:19 +00:00
David Roberts	cbe063a074	[ML] Introduce a "starting" datafeed state for lazy jobs (#53918 ) It is possible for ML jobs to open lazily if the "allow_lazy_open" option in the job config is set to true. Such jobs wait in the "opening" state until a node has sufficient capacity to run them. This commit fixes the bug that prevented datafeeds for jobs lazily waiting assignment from being started. The state of such datafeeds is "starting", and they can be stopped by the stop datafeed API while in this state with or without force. Fixes #53763	2020-03-24 10:01:13 +00:00
István Zoltán Szabó	8279f82dea	[DOCS] Fixes typo in start datafeed API docs. (#53811 )	2020-03-19 17:55:26 +01:00
István Zoltán Szabó	57321124ea	[DOCS] Changes seconds to milliseconds since the Epoch in AD docs. (#53797 )	2020-03-19 15:40:53 +01:00
Tom Veasey	58340c2dbe	[ML] Adds the class_assignment_objective parameter to classification (#52763 ) Adds a new parameter for classification that enables choosing whether to assign labels to maximise accuracy or to maximise the minimum class recall. Fixes #52427.	2020-03-12 18:39:29 +00:00
István Zoltán Szabó	77ec60baa0	[DOCS] Adds a warning about reindexing docs with the same ID to the PUT DFA docs. (#53490 )	2020-03-12 18:00:36 +01:00
Benjamin Trent	4e1f029b04	[ML][Inference] adds new default_field_map field to trained models (#53294 ) Adds a new `default_field_map` field to trained model config objects. This allows the model creator to supply field map if it knows that there should be some map for inference to work directly against the training data. The use case internally is having analytics jobs supply a field mapping for multi-field fields. This allows us to use the model "out of the box" on data where we trained on `foo.keyword` but the `_source` only references `foo`.	2020-03-11 12:23:56 -04:00
Dimitris Athanasiou	5a32f50d18	[ML] Rename data frame analytics maximum_number_trees to max_trees (#53300 ) Deprecates `maximum_number_trees` parameter of classification and regression and replaces it with `max_trees`.	2020-03-11 10:33:53 +02:00
István Zoltán Szabó	54b66d3385	[DOCS] Makes the description clearer on how to use aggregations in an anomaly detection job (#53103 ) Co-authored-by: lcawl <lcawley@elastic.co>	2020-03-09 09:48:23 +01:00
István Zoltán Szabó	08fcc0b02f	[DOCS] Adds deleting flag to the GET job stats API docs (#53223 )	2020-03-06 16:03:09 +01:00
István Zoltán Szabó	870e1891d9	[DOCS] Makes the naming convention of the DFA response objects coherent (#53172 )	2020-03-05 16:25:43 +01:00
István Zoltán Szabó	d7fb6416dd	[DOCS] Expands GET DFA stat API docs with response objects. (#53107 )	2020-03-05 15:30:30 +01:00
Lisa Cawley	7004216455	[DOCS] Adds link in datafeed indices_options (#53067 )	2020-03-03 10:28:54 -08:00
István Zoltán Szabó	24fe7e5899	[DOCS] Adds response body documentation to GET inference API (#53050 )	2020-03-03 16:25:24 +01:00
Lisa Cawley	b6534834f9	[DOCS] Adds cat anomaly detectors API (#52866 )	2020-02-28 12:15:21 -08:00

1 2 3 4 5

242 Commits