pgbackrest

mirror of https://github.com/pgbackrest/pgbackrest.git synced 2025-07-03 00:26:59 +02:00

Author	SHA1	Message	Date
David Steele	5dba0d6e9b	Set option-archive-copy flag in backup.manifest to false when offline. In offline mode the pg_wal directory is copied, but that is not the same as archive-copy, which copies the exact set of WAL required from the archive. This flag is purely for informational purposes so there is no live bug here, but the prior behavior was certainly misleading.	2022-04-05 18:42:19 -04:00
David Steele	14016a86e7	Check that sha1 checksum is not empty in manifestFileUpdate(). The manifest test module was setting a blank value here and causing a stack overflow because memcpy() is used instead of strcpy(). This was really just a test issue but add an assert just in case the same were to happen in production code. Also update a bogus checksum in the integration tests to the correct length to avoid running afoul of the assert. Found with -fsanitize=address.	2022-03-24 13:13:35 -06:00
David Steele	7afaac0a3d	Allow repo-hardlink option to be changed after full backup. This rule was added because there were not sufficient tests to demonstrate that the repo-hardlink option could be changed in a backup set. Remove the restriction and add/update tests to show that it works. This is necessary now because bundling requires that hardlinking be disabled. Rather than add code complexity, it seems better just to address this limitation.	2022-03-22 08:35:34 -06:00
David Steele	2c96327e65	Remove extraneous double spaces in code and comments.	2022-03-15 17:55:48 -06:00
David Steele	9eec98c613	Retry on page checksum validation failure during backup. Rather than attempting to filter page checksum failures by LSN, just retry when there is a page checksum failure. If the page has not changed since the last read report it as an error. If the page has changed, then PostgreSQL must be modifying the page so we can ignore the error because a full page write (and possibly updates) will be in the WAL. Also remove tests made redundant by the test merge in `b4897077`.	2022-02-23 12:05:53 -06:00
David Steele	34d649579e	Bundle files in the repository during backup. Bundle (combine) smaller files during backup to reduce the number of files written to the repository (enable with --bundle). Reducing the number of files is a benefit on all file systems, but especially so on object stores such as S3 that have a high file creation cost. Another benefit is that zero-length files are only stored as metadata in the manifest. Files are batched up to bundle-size and then compressed/encrypted individually and stored sequentially in the bundle. The bundle id and offset of each file is stored in the manifest so files can be retrieved randomly without needing to read the entire bundle. Files are ordered by timestamp descending when being assigned to bundles to reduce the amount of random access that needs to be done. The idea is that bundles with older files can be read in their entirety on restore and only bundles with newer files will get fragmented. Bundles are a custom format with metadata stored in the manifest. Tar was considered but it is too limited a format, the major issue being that the size of the file must be known in advance and that is very contrary to how pgBackRest works, especially once we introduce page-level incremental backups. Bundles are stored numbered in the bundle directory. Some files may still end up in pg_data if they are added after the backup is complete. backup_label is an example. Currently, only the backup command works in batches. The restore and verify commands use the offsets to pull individual files out of the bundle. It seems better to finalize how this is going to work before optimizing the other commands. Even as is, this is a major step forward, and all commands function with bundling. One caveat: resume is currently not supported when bundle is enabled.	2022-02-14 13:24:14 -06:00
David Steele	b0db4b8ff0	Simplify base path mode in mock/all integration tests. Change the mode back to 0700 earlier to reduce churn in the expect logs. This will be especially important in a future commit that gets the defaults exclusively from the base path.	2022-01-21 08:52:51 -05:00
David Steele	16559d9e42	Use the PG_FILE_POSTMTRPID constant where appropriate. Do the same in Perl with the MANIFEST_FILE_POSTMTRPID constant.	2022-01-20 08:41:05 -05:00
David Steele	bb4b30ddd3	Remove support for PostgreSQL 8.3/8.4. There is no evidence that users need 8.3/8.4 anymore but it does cost us in terms of development and testing, especially now that we have a number of new backup/restore features planned. It seems to make sense to remove this support now. If there are users who need to use/migrate from these versions they can use an older version of pgBackRest.	2022-01-06 15:34:04 -05:00
David Steele	43cfa9cef7	Revive archive performance test. This test was lost due to a syntax issue in `a58635ac`. Update the test to use system() to better mimic what postgres does and add logging so pgBackRest timing can be determined.	2021-11-10 12:14:41 -05:00
David Steele	ccc255d3e0	Add TLS Server. The TLS server is an alternative to using SSH for protocol connections to remote hosts. This command is currently experimental and intended only for trial and testing. As such, the new commands and options will not show up in the command-line help unless directly requested.	2021-10-18 14:32:41 -04:00
David Steele	01b20724da	Rename PostgreSQL pid file constants and tests.	2021-10-13 19:36:59 -04:00
David Steele	02b06aa495	Increase max index allowed for pg/repo options to 256. The prior limitations were based on using getopt_long() to parse command-line options, which required a static list of allowed options. Setting index max too high bloated the binary unacceptably. `45a4e80` replaced the functionality of getopt_long() but the static list remained. Improve cfgParseOption() to use available option data and remove the need for a static list. This also allows the option deprecations to be represented more compactly. Index max is still capped at 256 because a large enough index could cause parseOptionIdxValue() to run out of memory since it allocates a static list based on the highest index found. If that function were improved with a map of found index values then index max could be set to UINT64_MAX. Note that deprecations no longer set an index max or define whether reset is valid. These were space-saving measures which are no longer required. This means that indexed deprecated options will also be valid up to 256 and always allow reset, but it doesn't seem worth additional code to limit this behavior. cfgParseOptionId() is no longer needed because calling cfgParseOption() with .ignoreMissingIndex = true duplicates the functionality of cfgParseOptionId(). This leads to some simplification in the help code.	2021-08-31 12:09:50 -04:00
David Steele	eb98b8d2db	Fix typo.	2021-07-21 13:19:09 -04:00
David Steele	2452c4d5a4	Add PostgreSQL 14 support. There are no code changes from PostgreSQL 13 so simply add the new version. Add CATALOG_VERSION_NO_MAX to allow the catalog version to "float" during the PostgreSQL beta/rc period so new pgBackRest versions are not required when the catalog version changes. Update the integration tests to handle new PostgreSQL startup messages.	2021-05-24 17:17:03 -04:00
David Steele	01b8e2258f	Improve archive-push command fault tolerance. `3b8f0ef` missed some cases that could cause archive-push to fail: * Checking archive info. * Checking to see if a WAL segment already exists. These cases are now handled so archive-push can succeed on any valid repos.	2021-03-25 12:54:49 -04:00
Cynthia Shang	31c7824a4d	Allow stanza-* commands to be run remotely. The stanza-create, stanza-upgrade and stanza-delete were required to be run on the repository host. When there was only one repository allowed this was not a problem. However, with the introduction of multiple repository support, this becomes more of a burden to the user, therefore the stanza-create, stanza-upgrade and stanza-delete commands have been improved to allow for them to be run remotely.	2021-03-10 08:10:46 -05:00
David Steele	1dbb3bf50b	Multiple repository support. Up to four repositories may be configured. A potential benefit is the ability to have a local repository for fast restores and a remote repository for redundancy. Some commands, e.g. stanza-create/stanza-update, will automatically work with all configured repositories while others, e.g. stanza-delete, will require a repository to be specified using the repo option. See the command reference for details on which commands require the repository to be specified. Note that the repo option is not required when only repo1 is configured in order to maintain backward compatibility. However, the repo option is required when a single repo is configured as, e.g. repo2. This is to prevent command breakage if a new repository is added later. The archive-push command will always push WAL to the archive in all configured repositories but backups will need to be scheduled individually for each repository. In many cases this is desirable since backup types and retention will vary by repository. Likewise, restores must specify a repository. It is generally better to specify a repository for restores that has low latency/cost even if that means more recovery time. Only restore testing can determine which repository will be most efficient. For single repository configurations there should be no change in behavior.	2021-03-08 13:31:13 -05:00
David Steele	088662d986	GCS support for repository storage. GCS and GCS-compatible object stores can now be used for repository storage.	2021-03-05 12:13:51 -05:00
David Steele	bec3e20b2c	Add archive-get command multi-repo support. Repositories will be searched in order for the requested archive file. Errors will be reported as warnings as long as a valid copy of the archive file is found.	2021-02-23 15:34:28 -05:00
Cynthia Shang	f32eb9b94e	Partial multi-repository implementation. Multi-repository implementations for the archive-push, check, info, stanza-create, stanza-upgrade, and stanza-delete commands. Multi-repo configuration is disabled so there should be no behavioral changes between these commands and their current single-repo implementations. Multi-repo documentation and integration tests are still in the multi-repo development branch. All unit tests work as multi-repo since they are able to bypass the configuration restrictions.	2021-01-21 15:21:50 -05:00
David Steele	6e7a3eb383	Remove archive-timeout from test in mock/archive. No timeout is expected here but the small timeout prevents errors from being thrown. This is not a bug since the error would be thrown on the next archive-get call but it does make the tests harder to debug when there is an error. It is not clear why there was a timeout here at all. It is likely cruft from a prior test or a copy/paste error.	2021-01-05 18:11:28 -05:00
David Steele	ec9f23d31f	Remove CentOS 6 from tests and documentation. CentOS6 EOL'd and the mirrors were swiftly deleted, leading to failures in tests and documentation. Remove CentOS 6 for now to get builds going again with the intention to replace it in the near future with CentOS 8.	2020-12-02 16:23:05 -05:00
David Steele	2d38d2fc82	Reset additional options in real/all integration test. Currently indexes above 1 do not have dependencies checked, so this doesn't error. In a future commit we will enable those checks and this will error if it is not fixed.	2020-10-19 17:06:52 -04:00
David Steele	24d2c5b277	Remove real/all integration tests now covered by unit tests. Remove all check and stanza-* tests except for the ones that are intended to succeed. The successful tests show that the queries run with expected results against each version of PG which should also validate queries for the failure tests in the unit tests. Also remove the tests for --no-online backups since they don't require a database and are well tested in the unit tests.	2020-07-16 13:57:14 -04:00
David Steele	88b0f6245d	Run non version specific real/tests on the expect version. There are a few non version specific tests that need to be run in integration because we can't get coverage in the unit tests. To save some time we'll only run those tests against the same version we use for expect testing.	2020-07-15 13:19:16 -04:00
Stefan Fercot	d3dd32a031	Add expire-auto option. This allows automatic expiration after a successful backup to be disabled.	2020-07-14 08:12:25 -04:00
David Steele	682ac656f5	Fix restore --force acting like --force --delta. This caused restore to replace files based on timestamp and size rather than overwriting, which meant some files that should have been updated were left unchanged. Normal restore and restore --delta were not affected by this issue.	2020-07-06 15:03:24 -04:00
David Steele	3f4371d7a2	Azure support for repository storage. Azure and Azure-compatible object stores can now be used for repository storage. Currently only shared key authentication is supported but SAS will be added soon.	2020-07-02 16:24:34 -04:00
David Steele	a3e5e66f05	Simplify test matrix for real/all tests. Test matrices were previously simplified for the mock/* tests (e.g. `d4410611`, `d489eb87`) but not for real/all since the rules for which tests would run with which options was extremely complex. This only got more complex when new compression formats were added. Because the loop-generated matrix was so large, mosts tests were skipped for most option combinations following arcane logic which was nearly impossible to decipher even when reading the code, and completely impossible from the test.pl interface. As a consequence, important tests got excluded. For example, backup from standby was excluded for most versions of PostgreSQL because it was only run once per distro, against the latest version to be included in that distro. Simplify the tests by having a single run per PostgreSQL version and vary test parameters according to the capabilities of each version and the underlying distro. So, ZST testing is based on whether the distro supports ZST. Every test is run for each set of parameters based on the capabilities of the PostgreSQL version, e.g. backup from standby is not attempted on versions that don't support it. Note that since more tests are running the overall time to run the mock/all tests has increased by about 20-25%. Some time may be saved my removing tests that are adequately covered by unit tests but that should the subject of another commit. Another option would be to limit some non version-specific tests to a single, well defined version of PostgreSQL, .e.g the version that is run by expect tests, currently 9.6. The motivation for this refactor is that new storage drivers are coming and the loop-generated test matrix simply was not up to the task of adding them. The following is an example of the new test log (note longer runtime of each test): module=real, test=all, run=1, pg-version=10 (106.91s) module=real, test=all, run=1, pg-version=9.5 (151.09s) module=real, test=all, run=1, pg-version=9.2 (123.11s) module=real, test=all, run=1, pg-version=9.1 (129s) vs. the old test log (sub-second tests were skipped entirely): module=real, test=all, run=2, pg-version=10 (0.31s) module=real, test=all, run=3, pg-version=10 (0.26s) module=real, test=all, run=4, pg-version=10 (60.39s) module=real, test=all, run=1, pg-version=10 (69.12s) module=real, test=all, run=6, pg-version=10 (34s) module=real, test=all, run=5, pg-version=10 (42.75s) module=real, test=all, run=2, pg-version=9.5 (0.21s) module=real, test=all, run=3, pg-version=9.5 (0.21s) module=real, test=all, run=4, pg-version=9.5 (0.21s) module=real, test=all, run=5, pg-version=9.5 (0.26s) module=real, test=all, run=6, pg-version=9.5 (0.21s) module=real, test=all, run=1, pg-version=9.2 (72.78s) module=real, test=all, run=2, pg-version=9.2 (0.26s) module=real, test=all, run=3, pg-version=9.2 (0.31s) module=real, test=all, run=4, pg-version=9.2 (0.21s) module=real, test=all, run=5, pg-version=9.2 (0.21s) module=real, test=all, run=6, pg-version=9.2 (0.21s) module=real, test=all, run=1, pg-version=9.5 (88.41s) module=real, test=all, run=2, pg-version=9.1 (0.21s) module=real, test=all, run=3, pg-version=9.1 (0.26s) module=real, test=all, run=4, pg-version=9.1 (0.21s) module=real, test=all, run=5, pg-version=9.1 (0.31s) module=real, test=all, run=6, pg-version=9.1 (0.26s) module=real, test=all, run=1, pg-version=9.1 (72.4s)	2020-06-23 13:44:29 -04:00
David Steele	3d74ec1190	Use PostgreSQL instead of postmaster where appropriate. Using postmaster in messages was not very helpful since users rarely interact directly with the postmaster. Using PostgreSQL instead seems clearer.	2020-06-17 15:14:59 -04:00
David Steele	0680cfc8dc	Rename most instances of master to primary in tests. This aligns better with general PostgreSQL usage and our own documentation (updated in `4bcef702`). Usage in the backup.manifest tests has not been updated since it might break the file format.	2020-06-16 14:06:38 -04:00
David Steele	b5dd14e6f3	Make storage type more generic in the integration tests. Rather than bS3 use strStorage which can indicate more than two storage types. For the moment there are still only two storage types but this change is required before more can be added.	2020-05-12 18:55:20 -04:00
Stephen Frost	a021c9fe05	Add bzip2 compression support. bzip2 is a widely available, high-quality data compressor. It typically compresses files to within 10% to 15% of the best available techniques (the PPM family of statistical compressors), while being around twice as fast at compression and six times faster at decompression. bzip2 is currently available on all supported platforms.	2020-05-05 16:49:01 -04:00
David Steele	47aa765375	Add Zstandard compression support. Zstandard is a fast lossless compression algorithm targeting real-time compression scenarios at zlib-level and better compression ratios. It's backed by a very fast entropy stage, provided by Huff0 and FSE library. Zstandard version >= 1.0 is required, which is generally only available on newer distributions.	2020-05-04 15:25:27 -04:00
Cynthia Shang	c5241e5007	Expire WAL archive only when repo-retention-archive threshold is met. Previously when retention-archive was set (either by the user or by default), archives prior to the archive-start of the oldest remaining full backup (after backup expiration occurred) would be expired even though the retention-archive threshold had not been met. For example, if there were 1 full backup remaining after backup expiration and the retention-archive was set to 2 and retention-archive-type=full, then archives prior to the archive-start of the remaining full backup would still be removed even though retention-archive required 2 full backups remaining before archives should be expired. The thought was to keep the archive directory clean and since the full backup did not require prior archives, it was safe to delete them. However, this has caused problems for some users in the past (because they needed the WAL for other purposes) and with the new adhoc and time-based retention features, it was decided that the archives should remain until the threshold was met. The archives will eventually be removed and if having them causes space issues, the expire command and the retention-archive can always be run and adjusted.	2020-04-29 08:06:49 -04:00
Cynthia Shang	1c1a710460	Add --set option to the expire command. The specified backup set (i.e. the backup label provided and all of its dependent backups, if any) will be expired regardless of backup retention rules except that at least one full backup must remain in the repository.	2020-04-27 14:00:36 -04:00
David Steele	8af0462c5d	Fix race condition in real/all integration tests. If the tests are running quickly then the time target might end up the same as the end time of the prior full backup. That means restore auto-select will not pick it as a candidate and restore the last backup instead causing the restore compare to fail. So, sleep one second.	2020-03-26 15:30:59 -04:00
David Steele	4a5bd002c0	Move pgBackRest::Version module to pgBackRestDoc::ProjectInfo. The primary source for project info is now src/version.h. The pgBackRestDoc::ProjectInfo module loads the project info from src/version.h at runtime so there is no need to update it.	2020-03-10 17:57:02 -04:00
David Steele	731b862e6f	Rename BackRestDoc Perl module to pgBackRestDoc. This is consistent with the way BackRest and BackRest test were renamed way back in `18fd2523`. More modules will be moving to pgBackRestDoc soon so renaming now reduces churn later.	2020-03-10 15:41:56 -04:00
David Steele	36d4ab9bff	Move Perl modules out of lib directory. This directory was once the home of the production Perl code but since `f0ef73db` this is no longer true. Move the modules to test in most cases, except where the module is expected to be useful for the doc engine beyond the expected lifetime of the Perl test code (about a year if all goes well). The exception is pgBackRest::Version which requires more work to migrate since it is used to track pgBackRest versions.	2020-03-10 15:12:44 -04:00
David Steele	c279a00279	Add lz4 compression support. LZ4 compresses data faster than gzip but at a lower ratio. This can be a good tradeoff in certain scenarios. Note that setting compress-type=lz4 will make new backups and archive incompatible (unrestorable) with prior versions of pgBackRest.	2020-03-10 14:45:27 -04:00
David Steele	79cfd3aebf	Remove LibC. This was the interface between Perl and C introduced in `36a5349b` but since `f0ef73db` has only been used by the Perl integration tests. This is expensive code to maintain just for testing. The main dependency was the interface to storage, no matter where it was located, e.g. S3. Replace this with the new-introduced repo commands (`d3c83453`) that allow access to repo storage via the command line. The other dependency was on various cfgOption* functions and CFGOPT_ constants that were convenient but not necessary. Replace these with hard-coded strings in most places and create new constants for commonly used values. Remove all auto-generated Perl code. This means that the error list will no longer be maintained automatically so copy used errors to Common::Exception.pm. This file will need to be maintained manually going forward but there is not likely to be much churn as the Perl integration tests are being retired. Update test.pl and related code to remove LibC builds. Ding, dong, LibC is dead.	2020-03-09 17:41:59 -04:00
David Steele	3c4f91b319	Remove Perl unit tests made obsolete in `434cd832`. These were replaced by C unit tests but not all the unit test setup code was removed in the Perl module.	2020-03-09 13:35:26 -04:00
David Steele	438b957f9c	Add infrastructure for multiple compression type support. Add compress-type option and deprecate compress option. Since the compress option is boolean it won't work with multiple compression types. Add logic to cfgLoadUpdateOption() to update compress-type if it is not set directly. The compress option should no longer be referenced outside the cfgLoadUpdateOption() function. Add common/compress/helper module to contain interface functions that work with multiple compression types. Code outside this module should no longer call specific compression drivers, though it may be OK to reference a specific compression type using the new interface (e.g., saving backup history files in gz format). Unit tests only test compression using the gz format because other formats may not be available in all builds. It is the job of integration tests to exercise all compression types. Additional compression types will be added in future commits.	2020-03-06 14:41:03 -05:00
David Steele	02aa03d1a2	Remove obsolete methods in pgBackRest::Storage::Storage module. All the methods in this module will need to be implemented via the command-line in order to get rid of LibC, so the first step is to reduce the code in the module as much as possible. First remove storageDb() and use storageTest() instead. Then create storageTest() using pgBackRestTest::Common::Storage which has no dependencies on LibC. Now the only storage using the LibC interface is storageRepo(). Remove all link functions since those operations cannot be performed on a repo unless it is Posix, in which case the LibC interface is not needed. Same for owner(). Remove pathSync() because syncs are not required in the tests. No test data is reused after a crash. Path create/exists functions should never be explicitly performed on a repo so remove those. File exists can be implemented by calling info() instead. Remove encryption detection functions which were only used by Backup/Archive::Info reconstruct() which are now obsolete. Remove all filters except pgBackRest::Storage::Filter::CipherBlock since they are not being used. That also means there are no filters returning results so remove all the result code. Move hashSize() and pathAbsolute() into pgBackRest::Storage::Base where they can be shared between pgBackRest::Storage::Storage and pgBackRestTest::Common::Storage.	2020-03-06 14:10:09 -05:00
David Steele	00647c7109	Remove Perl Db module and LibC dependencies. This was mostly dead code except the DB_BACKUP_ADVISORY_LOCK constant, moved to the real/all test module, and the function that pulls info from pg_control, moved to ExpireEnvTest.pm.	2020-03-06 07:21:17 -05:00
David Steele	eb4347f20b	Use static checksums in mock/all integration tests. Using static values serves as a better cross-check against the page checksum code. The downside is that these checksums may not work with some big endian systems but in that case neither will the unit tests. We can also remove the page checksum interface from LibC which brings us one step closer to eliminating it.	2020-03-05 13:56:20 -05:00
David Steele	4ab8943ca8	Use PG_PAGE_SIZE_DEFAULT constant instead of pageSize variable. Page size is passed around a lot but in fact it can only have one value, PG_PAGE_SIZE_DEFAULT, which is checked when pg_control is loaded. There may be an argument for supporting multiple page sizes in the future but for now just use the constant to simplify the code. There is also a significant performance benefit. Because pageSize was being used in pageChecksumBlock() the main loop was neither unrolled nor vectorized (-funroll-loops -ftree-vectorize) as it is now with a constant loop boundary.	2020-03-05 09:14:27 -05:00
David Steele	91f321fb86	Rename old page*() functions to conform to new conventions. The general convention now is to prefix PostgreSQL functions with "pg".	2020-03-04 14:24:40 -05:00

1 2 3 4 5 ...

266 Commits