pgbackrest

mirror of https://github.com/pgbackrest/pgbackrest.git synced 2024-12-14 10:13:05 +02:00

Author	SHA1	Message	Date
David Steele	20782c88bc	PostgreSQL 15 support. PostgreSQL 15 drops support for exclusive backup and renames the start/stop backup commands. This is based on the pgdg-testing repo since beta1 has not been released yet, but it seems unlikely that breaking changes will be made at this point. beta1 should be tagged just before our next release so we'll retest before the release.	2022-05-04 11:55:59 -04:00
David Steele	692fe496bd	Remove dependency on pg_database.datlastsysoid. This column has been removed in PostgreSQL 15. Rather than add a lot of special handling, it seems better just to update all versions to not depend on this column. Add centralized functions to identify the type of database (i.e. system or user) by name and use FirstNormalObjectId when a name is not available. The new query in the db module will still return the prior result for PostgreSQL <= 15, which will be stored in the manifest. This is important to preserve behavior when downgrading pgBackRest. There are no concerns here for PostgreSQL 15 since older versions of pgBackRest won't be able to restore backups for PostgreSQL 15 anyway.	2022-05-04 08:22:45 -04:00
David Steele	9a271e925c	Fix error thrown from FINALLY() causing an infinite loop. Any error thrown resets execution to the last setjmp(), which means that parts of the try block need to make sure they don't get run again. FINALLY() was not doing this so if it threw an error it would end up back in the FINALLY() block, where the error would likely be thrown again, causing an infinite loop. Fix this by tracking the state of FINALLY() and only running it once. This requires cleaning the error stack like CATCH*() and clearing the error like TRY_END() depending on the order of execution.	2022-05-03 14:34:05 -04:00
David Steele	b89c568b5f	Fix obsolete variable naming.	2022-05-03 10:50:48 -04:00
David Steele	9629908694	Error on all lock failures except another process holding the lock. The archive-get/archive-push commands would not error for, .e.g permissions errors, when attempting to get a lock before launching the async process. Since the async process was not launched there would be no error status file and the user would get a generic failure message. Also, there would be no async log. Refactor lockAcquireFile() to throw an error when failOnNoLock = false unless the file is locked by another process. This seems to be the original intent of this parameter and there may have been a mistake when porting from Perl. In any case it looks wrong enough to be considered a bug.	2022-05-03 10:13:32 -04:00
David Steele	eb435becb3	Exclude mem context name from production builds. The mem context name is used to produce clearer debug errors but it has no purpose in production builds. Also remove memContextName() and access the struct directly since the name is only used within the common/memContext module. Note that a few errors that were thrown in production builds (and required the name) are now only thrown in debug builds. In practice we have not seen these errors in production builds due to extensive coverage so it does not seem worth modifying the error to work without the context name. This saves some memory, which is worthwhile, but the goal is to refactor Strings and Variants to have their own mem contexts and this change will prevent them from using more memory than they are now, along with other changes that will be coming later.	2022-05-02 15:17:34 -04:00
David Steele	0055fa40fe	Add user:group to lock permission error. This will help debug permissions errors when the lock file cannot be created.	2022-05-02 09:45:57 -04:00
David Steele	03c71aa606	Add hint to check the log on archive-get/archive-push async error. If this error is thrown rather than a specific error returned from the async process, it means the async process is unable to write the status files for some reason and the only way to get the error is out of the async log. This hint includes the exact async log path and name to make finding errors easier.	2022-05-02 08:49:13 -04:00
David Steele	4872a3f121	Improvements to test harness memory debugging. Only set -DDEBUG_MEM for the modules currently being tested rather than globally. Also run tests in a temp mem context. Running in the top context can confuse memory accounting when a new context is created in the top context.	2022-04-28 12:33:39 -04:00
David Steele	90f939b36f	Fix leaks in common/io unit test. These leaks make it harder to detect leaks in the core code, so fix them.	2022-04-28 12:31:59 -04:00
David Steele	8047e97e31	Fix leaked String and Variant in harnessPqScriptRun().	2022-04-28 12:17:33 -04:00
David Steele	083c93eaa3	Reuse Strings in iniLoad(). Reuse the section/key/value Strings by truncating them instead of creating a new one every time. Also add an error for empty sections. This function is only used for loading info files (not config files), which should never contain an empty section.	2022-04-28 10:11:15 -04:00
David Steele	bc46d4e37b	Add cvtZSubNTo*() functions. These functions allow conversion from substrings without needing to create a String or a temporary buffer. httpDateToTime() no longer requires a temp mem context. Also improve handling of month search to avoid an allocation. httpUriDecode() no longer requires a temp mem context. jsonReadStr() no longer requires a temp mem context. pgLsnFromWalSegment() no longer requires a temp mem context. pgVersionFromStr() no longer requires a temp mem context. Also do a bit of refactoring. storageGcsCvtTime() no longer leaks six Strings per call. storageS3CvtTime() no longer leaks six Strings per call.	2022-04-28 09:50:23 -04:00
David Steele	3f7c8bc923	Fix object allocations in incorrect mem context in execOpen(). Object variables were begin allocated in the calling context rather than the object context. This is not a live bug because Exec objects are currently created and opened in a long-lived context.	2022-04-26 10:15:47 -04:00
David Steele	41f9d69edc	Combine functions in the command/stanza module into one function. It is not clear why these were split out, but it probably had something to do with testing before storageList() could return NULL for an empty directory. Also remove the tests that depended on a boolean return, which are no longer needed for coverage.	2022-04-25 15:38:49 -04:00
David Steele	582c3dab4c	Add strLstAddSub() and strLstAddSubZ() functions. These help with readability and remove a cause of leaks.	2022-04-25 12:32:33 -04:00
David Steele	ff45f463cf	Use strLstAddZ() instead of strLstAdd() where possible. Using STRDEF() to convert the zero-terminated string to a String has no performance advantage but generates more code.	2022-04-25 11:58:30 -04:00
David Steele	7900660d3a	Add strLstNewFmt(). Simplifies adding a formatted string to a list and removes a common cause of leaks.	2022-04-25 11:47:43 -04:00
David Steele	45c3f4d53c	Improve JSON handling. Previously read/writing JSON required parsing/render via a variant, which add many more memory allocations and loops. Instead allow JSON to be read/written serially to improve performance and simplify the code. This also allows us to get rid of many String and Variant constant which are no longer required. The goal is to be able to read/write very large (e.g. gigabyte manifest) JSON structures, which would not be practical with the current code. Note that external JSON (GCS, S3, etc) is still handled using variants. Converting these will require more consideration about key ordering since it cannot be guaranteed as in our own formats.	2022-04-25 09:06:26 -04:00
David Steele	1e2b545ba4	Require type for FUNCTION_TEST_RETURN*() macros. This allows code to run after the return type has been generated in the case where it is an expression. No new functionality here yet, but this will be used by a future commit that audits memory usage.	2022-04-24 19:19:46 -04:00
David Steele	a2eee156b5	Fix instances where STRDEF() was used instead of STR(). In practice this didn't cause problems because the string buffer was still valid and strSize() was not being called.	2022-04-21 18:23:17 -04:00
David Steele	e18b70bf55	Allow RETURN() macros to accept struct initializers. Struct initializers look like multiple parameters in a macro so use __VA_ARGS__ to reconstruct them.	2022-04-21 07:45:59 -04:00
David Steele	ea4d73f375	Fix ordering of backup-lsn-stop field in command/restore unit test. All fields should be alphabetical. Currently the read code is tolerant of this, but that will not always be the case. Fields are always written alphabetically so this is just a test issue introduced by `d8d41321`.	2022-04-20 19:56:26 -04:00
David Steele	cb7a5f1ef3	Add JSON error when value does not parse in Ini object. If the JSON value fails to parse it is helpful to have the error message, at least for debugging.	2022-04-20 19:49:23 -04:00
David Steele	da6b4abc58	Handle missing archive start/stop in info/info backup unit test. This is not a very realistic case since archive start/stop are always written, but it appears in many other unit tests so it should also be tested here.	2022-04-20 19:41:28 -04:00
David Steele	d897bf1ec2	Add size to info/manifest unit test. This prevents the check from being order dependent.	2022-04-20 19:36:33 -04:00
David Steele	c304fafd45	Refactor PgClient to return results in Pack format. Packs support stronger typing than JSON and are more efficient. For the small result sets that we deal with efficiency is probably not very important, but this removes another place where we are using JSON instead of Pack. Push checking for result struct (e.g. single row) down into PgClient since it has easy access to this information rather than needing to parse the result set to find out. Refactor all code downstream that depends on PgClient results.	2022-04-20 08:36:53 -04:00
David Steele	cfd6c7ceb4	Use specific integer types in postgres/client and db unit tests. This will work better once we are able to transmit the results with stronger typing. Also remove int2 which was not being used.	2022-04-18 12:14:22 -04:00
David Steele	9751ddc4f8	Update postgres/client unit test to conform to current patterns. This includes adding test titles and using constants for query and error values that repeat.	2022-04-18 11:53:31 -04:00
David Steele	bc5f6fac34	Update postgres/client unit test for changes in libpq. There have been some behavioral changes in libpq which require changes to the test. Also update the instructions since it is now a bit easier to run against a real cluster.	2022-04-18 10:47:44 -04:00
David Steele	d103dd6238	Return stats as a JSON string rather than a KeyValue object. There is no need to process the stats so a KeyValue is overkill. Also remove the performance tests that check the stat totals since this is covered in the unit tests.	2022-04-14 20:34:42 -04:00
David Steele	e1ce731f8a	Add test for protocol greeting when a field is missing. A missing field and a NULL field are not exactly the same so it seems best to test both. Because of the way KeyValue objects work the error is the same, but that will not always be true.	2022-04-14 19:37:03 -04:00
David Steele	aeecd07ad8	Fix reported error line number when ini key length is zero. The line number was one less than it should have been, which could cause some confusion. Since this only affected ini files with JSON values, which are always written programmatically, there is almost zero chance this has ever been a problem in the field.	2022-04-14 18:29:54 -04:00
David Steele	fa40bcdc5c	Throw error when unable to read lock process. Previously the process id was skipped if it did not exist. Instead, throw an error and handle the errors in downstream code. This was probably ignored at some point to provide backward-compatibility, but that is no longer required, if it ever was.	2022-04-11 14:08:16 -04:00
David Steele	79b2041663	Add lockRead*() functions for reading locks from another process. Sometimes we need to read a lock from another process. This was done two different ways and in the case of cmdStop() was definitely hacky. Centralize the logic to make it easier to read the locks for another process. This will also make it easier to add new lock data.	2022-04-08 15:55:41 -04:00
Reid Thompson	aad7171940	Suppress existing WAL warning when archive-mode-check is disabled. When archive-mode-check is disabled and archive-push is running from multiple hosts, it is very likely that the file will already exist with the same checksum, so disable the warning. However, if the checksums do not match, an error will still be thrown.	2022-04-08 15:00:20 -04:00
David Steele	4f543a4d67	Handle NULL path in TEST_STORAGE_LIST when remove is specified. Using the path variable directly resulted in a path with (null) in it, which caused the remove to fail. The pathFull variable already exists for this purpose so use it.	2022-04-08 11:07:26 -04:00
David Steele	571dceefec	Add LENGTH_OF() macro. Determining the length of arrays that could be calculated at compile time was a bit piecemeal, with special macros used sometimes and with the math done directly other times. This macro makes the task easier, uses less space, and automatically adjusts when the type changes.	2022-04-07 19:00:15 -04:00
David Steele	8be11d32e4	Replace strCatFmt() with strCat()/strCatZ() where appropriate. Most of these looked like copy/paste from a prior required strCatFmt() call. There is no issue here since strCatFmt() works the same in these cases, but using strCat()/strCatZ() is more efficient.	2022-04-07 11:44:45 -04:00
David Steele	cff147a7d2	Add default for boolean options with unresolved dependencies. If a boolean option had an unresolved dependency then the value would be NULL, which meant the dependency would need to be checked in the code to avoid an error. For example, cfgOptionBool(cfgOptOnline) needed to be checked before it was safe to call cfgOptionBool(cfgOptArchiveCheck). Allow a default for boolean options when they are unresolved to simplify the code. This makes using the options easier and less prone to error. Not all boolean options get a dependency default in this commit, but more may be added in the future.	2022-04-06 14:45:51 -04:00
David Steele	5dba0d6e9b	Set option-archive-copy flag in backup.manifest to false when offline. In offline mode the pg_wal directory is copied, but that is not the same as archive-copy, which copies the exact set of WAL required from the archive. This flag is purely for informational purposes so there is no live bug here, but the prior behavior was certainly misleading.	2022-04-05 18:42:19 -04:00
David Steele	54b4187527	Show Docker output when building containers if --log-level=detail. This helps with debugging and monitoring container builds.	2022-04-05 13:14:42 -04:00
Reid Thompson	d8d4132118	Auto-select backup for restore command --type=lsn. For PITR with --type=lsn, attempt to auto-select the appropriate backup set based on the --target LSN provided. Pick the most recent backup where backup-lsn-stop is less than or equal to the provided LSN.	2022-04-05 11:59:12 -04:00
David Steele	f60ec5055a	Cleanup output to stderr in unit tests. The unit tests were ignoring stderr but nothing being output there was important. Now a test will fail if there is anything on stderr. This makes it easier to work with -fsanitize, which outputs to stderr.	2022-03-24 18:43:43 -06:00
David Steele	14016a86e7	Check that sha1 checksum is not empty in manifestFileUpdate(). The manifest test module was setting a blank value here and causing a stack overflow because memcpy() is used instead of strcpy(). This was really just a test issue but add an assert just in case the same were to happen in production code. Also update a bogus checksum in the integration tests to the correct length to avoid running afoul of the assert. Found with -fsanitize=address.	2022-03-24 13:13:35 -06:00
David Steele	75b26319ae	Use strNewZ() in cases where STRDEF() assignment goes out of scope. If a variable assigned with STRDEF() is referenced out of scope of the STRDEF() assignment then the value is undefined. Luckily most of the instances are in tests but there is one in the core code. It is not clear if this is a live bug or not but it certainly needs to be fixed. Found with -fsanitize=address.	2022-03-24 12:26:09 -06:00
David Steele	edf6c70baa	Prevent signed integer overflow in cfgParseSize(). If the value and multiplier were large enough then the return value could overflow unpredictably. Check the value to make sure it will not overflow with the current multiplier. It would be better to present an "out of range" error to the user rather than "is not valid" but it doesn't seem worth the effort since the error is extremely unlikely. Found with -fsanitize=undefined.	2022-03-24 11:00:51 -06:00
David Steele	ccbe2a1f70	Do not pass NULL to memcpy() in Buffer/String objects. glibc and others seem tolerant of this but the behavior is undefined. Found with -fsanitize=undefined.	2022-03-24 09:32:18 -06:00
David Steele	98792b1b0c	Do not pass NULL to bsearch()/qsort() in List object. glibc and others seem tolerant of this but the behavior is undefined. Found with -fsanitize=undefined.	2022-03-24 09:22:05 -06:00
David Steele	424008d293	Allow files that become zero-length after the backup manifest is built. It is possible that a file will be be truncated to zero-length after the backup manifest has been built. We could build logic into backupFile() to handle this case but it is hard to test well because of the race condition so tests would need to written directly against backupFile() and backupJobResult(). It hardly seems worth all that effort for a condition that occurs rarely, if ever. Instead just remove the manifest check and add tests to restore to make sure it handles bundled zero-length files correctly. Logging will show that the file was bundled so if it happens a lot (which seems very unlikely) then we can think about an alternate implementation.	2022-03-23 10:41:36 -06:00
David Steele	7afaac0a3d	Allow repo-hardlink option to be changed after full backup. This rule was added because there were not sufficient tests to demonstrate that the repo-hardlink option could be changed in a backup set. Remove the restriction and add/update tests to show that it works. This is necessary now because bundling requires that hardlinking be disabled. Rather than add code complexity, it seems better just to address this limitation.	2022-03-22 08:35:34 -06:00
Reid Thompson	5ae84d5e47	Improve path validation for repo-* commands. Check for invalid path in repo-* commands. Perform path validation and throw an error when appropriate. Path may not contain '//'. Strip trailing '/' from path. Absolute path must fall under repo path.	2022-03-22 07:50:26 -06:00
nunopi	21cef09dfd	Add AWS IMDSv2 support. IMDSv2 provides additional security to prevent instance metadata from being read by an attacker. All AWS instances should provide IMDSv2 but still fail back to IMDSv1 if the IMDSv2 token request fails. This is in case there are any services outside AWS that are emulating IMDSv1 but have not implemented IMDSv2.	2022-03-16 11:02:29 -06:00
David Steele	2c96327e65	Remove extraneous double spaces in code and comments.	2022-03-15 17:55:48 -06:00
David Steele	3f66f42ef9	Rename bundle-* options to repo-bundle-*. It seems best for these to be repo options so they can be configured per repo, rather than globally. All clarify usage for repo-bundle-size and repo-bundle-limit.	2022-03-14 17:49:52 -06:00
Reid Thompson	7c9208ba85	Improve error message for invalid repo-azure-key. Check that repo-azure-key is valid base64 when repo-azure-key-type = shared.	2022-03-11 10:10:02 -06:00
David Steele	0054677147	Add bundle logging to backup command. This was added to the restore command so add it to the backup command as well.	2022-03-09 15:34:15 -06:00
David Steele	dca6da86bf	Optimize restore command for file bundling. Since files are stored sequentially in a bundle, it is often possible to restore multiple files with a single read. Previously, each restored file required a separate read. Reducing the number of reads is particularly beneficial for object stores, but performance should benefit on any file system. Currently if there is a gap then a new read is required. In the future we might set a limit for how large a gap we'll skip without starting a new read.	2022-03-09 15:03:28 -06:00
Reid Thompson	f7ab002aa7	Improve stop command to honor stanza option. Improve the stop command, when force and stanza options are specified, to terminate only processes holding lock files for the given stanza. Prior to these changes, termination of all processes holding lock files regardless of stanza occurred.	2022-03-08 12:18:23 -06:00
David Steele	514137040e	Add limit parameter to ioCopyP(). Allows the number of bytes copied to be limited.	2022-03-08 08:23:31 -06:00
Reid Thompson	330e19900e	Increase precision of percent complete logging for backup and restore. For very large backups only getting an update per percent may not be often enough. Add hundredths to the percent complete logging to provide more timely information.	2022-03-06 13:01:24 -06:00
David Steele	8f23b46b4b	Replace percentage and size with a constant in restore test logs. Checking percentage and size in every test can cause quite a bit of churn when changes are made. Follow the example of the backup tests and replace percentage and size after the few tests to reduce churn.	2022-03-06 11:57:20 -06:00
David Steele	4d2fef1c37	Remove redundant restoreFile() test and improve coverage. These tests were written before the restore command was fully migrated to C so many of them have become redundant. In the cases were they still provide coverage, add tests to synthetic restores to replace them. In general, these higher level tests provide better coverage than poking at the restoreFile() function directly.	2022-03-06 11:48:22 -06:00
David Steele	5249b89a2e	v2.38: Minor Bug Fixes and Improvements IMPORTANT NOTE: Repository size reported by the info command is now entirely based on what pgBackRest has written to storage. Previously, in certain cases, pgBackRest could detect if additional compression was being applied by the storage but this is no longer supported. Bug Fixes: * Retry errors in S3 batch file delete. (Reviewed by Reid Thompson. Reported by Alex Richman.) * Allow case-insensitive matching of HTTP connection header values. (Reviewed by Reid Thompson. Reported by Rémi Vidier.) Features: * Add support for AWS S3 server-side encryption using KMS. (Contributed by Christoph Berg. Reviewed by David Steele, Tharindu Amila.) * Add archive-missing-retry option. (Reviewed by Stefan Fercot.) * Add backup type filter to info command. (Contributed by Stefan Fercot. Reviewed by David Steele.) Improvements: * Retry on page validation failure during backup. (Reviewed by Stephen Frost, David Christensen.) * Handle TLS servers that do not close connections gracefully. (Reviewed by Rémi Vidier, David Christensen, Stephen Frost.) * Add backup LSNs to info command output. (Contributed by Stefan Fercot. Reviewed by David Steele.) * Automatically strip trailing slashes for repo-ls paths. (Contributed by David Christensen. Reviewed by David Steele.) * Do not retry fatal errors. (Reviewed by Reid Thompson.) * Remove support for PostgreSQL 8.3/8.4. (Reviewed by Reid Thompson, Stefan Fercot.) * Remove logic that tried to determine additional file system compression. (Reviewed by Reid Thompson, Stefan Fercot.) Documentation Bug Fixes: * Move repo options in TLS documentation to the global section. (Reported by Anton Kurochkin.) * Remove unused backup-standby option from stanza commands. (Reported by Stefan Fercot.) * Fix typos in help and release notes. (Fixed by Daniel Gustafsson. Reviewed by David Steele.) Documentation Improvements: * Add aliveness check to systemd service configuration. (Suggested by Yogesh Sharma.) * Add FAQ explaining WAL archive suffix. (Contributed by Stefan Fercot. Reviewed by David Steele.) * Note that replications slots are not restored. (Contributed by Reid Thompson. Reviewed by David Steele, Stefan Fercot. Suggested by Christophe Courtois.)	2022-03-06 10:30:59 -06:00
David Steele	59a5373cf8	Handle TLS servers that do not close connections gracefully. Some TLS server implementations will simply close the socket rather than correctly closing the TLS connection. This causes problems when connection: close is specified with no content-length or chunked encoding and we are forced to read to EOF. It is hard to know if this is a real EOF or a network error. In cases where we can parse the content and (hopefully) ensure it is correct, allow the closed socket to serve as EOF. This is not ideal, but the change in `8e1807c` means that currently working servers with this issue will stop working after 2.35 is installed, which seems too risky.	2022-03-02 11:38:52 -06:00
David Steele	fb5051fde7	Use vagrant user in the Docker container. This is a bit of legacy from the current Vagrant environment used to do the release, but since it is not as easy to change the user in Vagrant, just make the Docker environment conform. This allows documentation to be built in a Vagrant environment (or any environment with the same user name) and to be deployed in a Docker environment.	2022-02-26 13:50:30 -06:00
David Steele	b33cabe08c	Allow case-insensitive matching of HTTP connection header values. The specification allows values for the connection header to be case-insensitive. See https://www.rfc-editor.org/rfc/rfc7230#section-6.1.	2022-02-25 10:51:40 -06:00
David Christensen	6320712323	Automatically strip trailing slashes for repo-ls paths. Trailing slashes in at least some of the repository storage types were preventing repo-ls from displaying any content (presumably due to storage-specific behavior). Since the path with the slash should be equivalent to the path without the slash, just remove it if provided by the user.	2022-02-23 13:53:02 -06:00
David Steele	53f1b25204	Improve validation of zero pages. Checking that pd_upper == 0 is not enough since this field may be corrupted. Still use pd_upper as a quick check, but when it is zero proceed to check the rest of the page to ensure it is also all zeroes.	2022-02-23 13:17:14 -06:00
David Steele	9eec98c613	Retry on page checksum validation failure during backup. Rather than attempting to filter page checksum failures by LSN, just retry when there is a page checksum failure. If the page has not changed since the last read report it as an error. If the page has changed, then PostgreSQL must be modifying the page so we can ignore the error because a full page write (and possibly updates) will be in the WAL. Also remove tests made redundant by the test merge in `b4897077`.	2022-02-23 12:05:53 -06:00
David Steele	67bdf07e69	Add XML to invalid XML error message. There have been cases where pgBackRest has failed on invalid XML but it is not possible to determine what was wrong with the XML. This will only work for XML up to about 8KiB (which is the error message limit) but it should work in most cases.	2022-02-23 10:26:39 -06:00
David Steele	10038db9c9	Add archive-missing-retry option. Retry a WAL segment that was previously reported as missing by the archive-get command. This prevents notifications in the spool path from a prior restore from being used and possibly causing a recovery failure if consistency has not been reached. Disabling this option allows PostgreSQL to more reliably recognize when the end of the WAL in the archive has been reached, which permits it to switch over to streaming from the primary. With retries enabled, a steady stream of WAL being archived will cause PostgreSQL to continue getting WAL from the archive rather than switch to streaming. When disabling this option it is important to ensure that the spool path for the stanza is empty. The restore command does this automatically if the spool path is configured at restore time. Otherwise, it is up to the user to ensure the spool path is empty.	2022-02-23 09:14:27 -06:00
David Steele	e6e1122dbc	Pass file by reference in manifestFileAdd(). Coverity complained that this pass by value was inefficient: CID 376402: Performance inefficiencies (PASS_BY_VALUE) Passing parameter file of type "ManifestFile" (size 136 bytes) by value. This was completely intentional since it gives us a copy of the struct that we can change without bothering the caller. However, updating fields is fine and may benefit the caller at some future data, and in any case does no harm now. And as usual it is easier not to fight with Coverity.	2022-02-20 16:45:07 -06:00
David Steele	b489707793	Move command/backup-common tests in the command/backup module. As much as possible it is better to get coverage with more realistic tests. Merging these modules will allow the page checksum code to be covered with real backups.	2022-02-18 17:50:05 -06:00
David Steele	efc09db7b9	Limit files that can be bundled. Limit which files can be added to bundles, which allows resume to work reasonably well. On resume, the bundles are removed and any remaining file is eligible to be to be resumed. Also reduce the bundle-size default to 20MiB. This is pretty arbitrary, but a smaller default seems better.	2022-02-17 07:25:12 -06:00
David Steele	34d649579e	Bundle files in the repository during backup. Bundle (combine) smaller files during backup to reduce the number of files written to the repository (enable with --bundle). Reducing the number of files is a benefit on all file systems, but especially so on object stores such as S3 that have a high file creation cost. Another benefit is that zero-length files are only stored as metadata in the manifest. Files are batched up to bundle-size and then compressed/encrypted individually and stored sequentially in the bundle. The bundle id and offset of each file is stored in the manifest so files can be retrieved randomly without needing to read the entire bundle. Files are ordered by timestamp descending when being assigned to bundles to reduce the amount of random access that needs to be done. The idea is that bundles with older files can be read in their entirety on restore and only bundles with newer files will get fragmented. Bundles are a custom format with metadata stored in the manifest. Tar was considered but it is too limited a format, the major issue being that the size of the file must be known in advance and that is very contrary to how pgBackRest works, especially once we introduce page-level incremental backups. Bundles are stored numbered in the bundle directory. Some files may still end up in pg_data if they are added after the backup is complete. backup_label is an example. Currently, only the backup command works in batches. The restore and verify commands use the offsets to pull individual files out of the bundle. It seems better to finalize how this is going to work before optimizing the other commands. Even as is, this is a major step forward, and all commands function with bundling. One caveat: resume is currently not supported when bundle is enabled.	2022-02-14 13:24:14 -06:00
David Steele	8046f06307	Do not retry fatal errors. There is some evidence that retrying fatal errors, especially out of memory errors, may cause lockups. It makes sense to report fatal errors as quickly as possible and bypass retries. This may or not fix the lockup issue but it is worth doing either way. For now, the only fatal errors will be AssertError and MemoryError.	2022-02-14 11:07:02 -06:00
David Steele	8d0cce66f8	Use normal error for protocol module error retry test. Asserts will not be retried in a future commit, so adjust this test now to use non-assert errors.	2022-02-13 15:19:31 -06:00
David Steele	8573a2df14	Improve protocol module error test for protocolClientFree(). Using an assert here was never ideal and won't work once we start handling fatal errors differently.	2022-02-13 15:11:59 -06:00
David Steele	551e5bc6f6	Retry errors in S3 batch file delete. If the entire batch failed it would be retried, but individual file errors were not retried. This could cause pgBackRest to terminate during expiration or when removing an unresumable backup. Rather than retry the entire batch, delete the errored files individually to take advantage of the HTTP retry rather than adding a new retry loop. These errors seem rare enough that it should not be a performance issue.	2022-02-11 08:11:39 -06:00
Stefan Fercot	b26097f8d8	Add backup type filter to info command. Support --type option in the info command to display only a specific backup type.	2022-02-09 10:18:39 -06:00
David Steele	cb630ffe3b	Remove logic that tried to determine additional file system compression. In theory, the additional stat() call after a file has been copied to the repo can determine if additional compression has been applied by the file system. However, it has been a very long time since we tested this in practice. There are currently no unit tests that accurately test this feature since it requires a compressed file system like ZFS to work, which never seemed worth the extra cost. It can also add a lot of time to backups if there are a large quantity of small files. In addition, it stands as a blocker for combining files for small file support since it is no longer possible to get per-file sizes from the viewpoint of the file system. There are several ways this could be reworked but none of them are easy while at the same time maintaining current info functionality. It doesn't seem worth keeping an untested feature that will only work in some special cases (if it still works) when it is blocking development.	2022-02-09 09:32:23 -06:00
David Steele	b1da4e84e8	Revert Minio to prior release. The most recent release of Minio has broken CI builds but there is no logging to indicate what is wrong. For now, just use the prior release to get CI builds working again. This kind if breakage is not uncommon for Minio but they usually resolve it in the next release.	2022-02-02 14:39:39 -06:00
David Steele	9b2f10dbb4	Refactor lock code. Update lock code to use standard common/io functions and module patterns. This module was developed before the common/io module existed and our patterns had stabilized.	2022-01-31 16:48:28 -06:00
David Steele	22734eb376	Add ioBufferReadNewOpen() and ioBufferWriteNewOpen(). These are convenience functions to make the code a bit more compact where possible.	2022-01-31 10:03:56 -06:00
David Steele	cf5b3a302f	Fix language in rh7 test container for aarch64. The /etc/profile.d/lang.sh script was causing issues but it does not exist on amd64, so it seems the easiest thing was to remove it. Fix how 32-bit VMs are determined now that another 64-bit architecture has been added. And remove some obsolete VM hashes.	2022-01-26 13:22:31 -06:00
David Steele	e4df5b7d38	Simplify manifest file defaults. Previously manifest load required two passes through the file list, one to load the data and one to set the defaults. This required each file to be packed twice. Instead simply note that the file value is default and then set the file defaults when they are loaded from the manifest. This is made possible by the different internal/external representations for files so the same method cannot be applied to paths and links. This change seems to resolve the performance issues noted in `61ce586` but there is no obvious reason why.	2022-01-24 15:21:07 -06:00
David Steele	61ce58692f	Pack manifest file structs to save memory. Manifests with a very large number of files can use a considerable amount of memory. There are a lot of zeroes in the data so it can be stored more efficiently by using base-128 varint encoding for the integers and storing the strings in the same allocation. The downside is that the data needs to be unpacked in order to be used, but in most cases this seems fast enough (about 10% slower than before) except for saving the manifest, which is 10% slower up to 10 million files and then gets about 5x slower by 100 million (two minutes on my M1 Mac). Profiling does not show this slowdown so I wonder if this is related to the change in memory layout. Curiously, the function that increased most was jsonFromStrInternal(), which was not modified. That gives more weight to the idea that there is some kind of memory issue going on here and one hopes that servers would be less affected. Either way, they largest use cases we have seen are for about 6 million files so if we can improve that case I believe we will be better off. Further analysis showed that most of the time was taken up writing the size and timestamp fields, which makes almost no sense. The same amount of time was used if they were hard-coded to 0, which points to some odd memory issue on the M1 architecture. This change has been planned for a while, but the particular impetus at this time is that small file support requires additional fields that would increase manifest memory usage by about 20%, even if the feature is not used. Note that the Pack code has been updated to use the new varint encoder, but the decoder remains separate because it needs to fetch one byte at a time.	2022-01-21 17:05:07 -05:00
David Steele	4a73a02863	Simplify manifest defaults. Manifest defaults for user, group, and mode were previously generated by scanning the data to find the most common values. This was very accurate but slow and complicated. It could also lead to surprising changes in the manifest when a default value suddenly changed. Instead, use the $PGDATA path to generate defaults. In the vast majority of cases the same user/group should own all the path/files and the default file mode is easily derived from the path mode. There may be some edge cases where this generates larger manifests, but in general it reduces time and complexity when saving the manifest. Remove the MCV code since it is longer longer used.	2022-01-21 15:22:48 -05:00
David Steele	b0db4b8ff0	Simplify base path mode in mock/all integration tests. Change the mode back to 0700 earlier to reduce churn in the expect logs. This will be especially important in a future commit that gets the defaults exclusively from the base path.	2022-01-21 08:52:51 -05:00
David Steele	8c062e1af8	Remove primary flag from manifest. This flag was only being used by the backup command after manifestNewBuild() and had no other uses. There was a time when it was important for integration testing but the unit tests now fulfill this role. Since backup is the only code concerned with the primary flag, move the code into the backup module. We don't have any cross-version testing but this change was tested manually with the most recent version of pgBackRest to make sure it was tolerant of the missing primary info. When an older version of pgBackRest loads a newer manifest the primary flag will always be set to false, which is fine since it is not used.	2022-01-20 14:01:10 -05:00
David Steele	16559d9e42	Use the PG_FILE_POSTMTRPID constant where appropriate. Do the same in Perl with the MANIFEST_FILE_POSTMTRPID constant.	2022-01-20 08:41:05 -05:00
David Steele	e21ba7c92b	Remove extra spaces.	2022-01-18 17:40:53 -05:00
David Steele	b791f1c82f	Implement restore ownership without updating manifest internals. Updating the manifest this way was not a great idea because it broke abstraction for the object. This meant certain changes to the interface and internals were not possible because the code was modifying internal manifest data. Instead track the user replacements entirely in the restore module. This also has the benefit of eliminating a pass over the manifest path/file/link lists.	2022-01-15 14:33:38 -05:00
Christoph Berg	3097acd73a	Add support for AWS S3 server-side encryption using KMS. AWS S3 integrates with AWS Key Management Service (AWS KMS) to provide server side encryption of S3 objects. This integration protects objects under encryption keys that never leave AWS KMS unencrypted.	2022-01-13 08:46:14 -05:00
David Steele	a79034ae2f	Add read range to all storage drivers. The range feature allows reading out an arbitrary chunk of a file and will be important for efficient small file support. Now that all drivers are required to support ranges remove the storageFeatureLimitRead feature flag that was implemented only by the Posix driver.	2022-01-11 14:42:53 -05:00
David Steele	2cddbbdee0	Remove obsolete cfgOptionHostPort()/cfgOptionIdxHostPort(). These functions were made obsolete by the refactor in `6a124584`.	2022-01-10 17:20:48 -05:00
David Steele	aeecb500f5	Improve implementation of cfgOptionIdxName(). Cache option names after they are generated rather than regenerating them each time.	2022-01-10 14:47:29 -05:00
David Steele	aced5d47ed	Replace cfgOptionGroupIdxToKey() with cfgOptionGroupName(). Do the replacement anywhere cfgOptionGroupIdxToKey() is being used to construct a group name in a message. cfgOptionGroupName() is better for this case since it also includes the name of the group so that it does not need to be repeated in each message.	2022-01-10 09:10:06 -05:00
David Steele	e4b48eb430	Fix inconsistent group display names in messages. In other instances there are no dashes, e.g. repo1 or pg1. Make these messages match.	2022-01-09 19:43:44 -05:00

1 2 3 4 5 ...

2218 Commits