pgbackrest

mirror of https://github.com/pgbackrest/pgbackrest.git synced 2024-12-14 10:13:05 +02:00

Author	SHA1	Message	Date
David Steele	bd2ba802db	Check that clusters are alive and correctly configured during a backup. Fail the backup if a cluster stops or the standby is promoted. Previously, shutting down the primary would cause an error but it was not detected until the end of the backup. Now the error will happen sooner and a promotion on the standby will also cause an error.	2021-12-08 10:16:41 -05:00
David Steele	7b3ea883c7	Add SIGTERM and SIGHUP handling to TLS server. SIGHUP allows the configuration to be reloaded. Note that the configuration will not be updated in child processes that have already started. SIGTERM terminates the server process gracefully and sends SIGTERM to all child processes. This also gives the tests an easy way to stop the server.	2021-12-07 18:18:43 -05:00
David Steele	49145d72ba	Add timeline and checkpoint checks to backup. Add the following checks: * Checkpoint is updated in pg_control after pg_start_backup(). This helps ensure that PostgreSQL and pgBackRest have a consistent view of the storage and that PGDATA paths match. * Timeline of backup start WAL file matches pg_control. Hard to see how this one could get hit, but we have the power... * Standby is on the same timeline as the primary. If not, this standby is not following the primary. * Last standby checkpoint is not greater than the backup checkpoint. If so, this standby is not following the primary. This also requires some additional plumbing to read/write timeline/checkpoint from pg_control and parse timelines from WAL filenames. There were some changes in the backup tests caused by the fact that pg_control now has different contents for each backup. The check to ensure that the required checkpoint was reached on the standby should also be updated to use pg_control (it currently uses pg_control_checkpoint()), but that requires non-trivial changes to the test harness and will need to wait.	2021-12-07 09:21:07 -05:00
Reid Thompson	dcb4f09d83	Revert changes to backupFilePut() made in `1e77fc3d`. These changes were made obsolete by `a3d7a23a`.	2021-11-23 09:37:12 -05:00
Reid Thompson	a3d7a23a9d	Use infoBackupDataByLabel() to log backup size. Eliminate summing and passing of copied files sizes for logging backup size. Instead, utilize infoBackupDataByLabel() to pull the backup size for the log message.	2021-11-22 12:52:37 -05:00
Reid Thompson	1a0560d363	Allow y/n arguments for boolean command-line options. This allows boolean boolean command-line options to work like their config file equivalents. At least for now this behavior will remain undocumented since all examples in the documentation will continue to use the standard syntax. The idea is that it will "just work" when options are copied out of config files rather than generating an error.	2021-11-19 12:22:09 -05:00
David Steele	2d963ce947	Rename server-start command to server.	2021-11-18 17:23:11 -05:00
David Steele	1f14f45dfb	Check archive immediately after backup start. Previously the archive was only checked at the end of the backup to ensure all WAL required to make the backup consistent was present. The problem was that if archiving was not functioning then the backup had to complete before the user found out, which could be a while if the database was large enough. Add an archive check immediately after backup start so failures are reported earlier. The trick is to determine which WAL to check. If the repo is new there may not be any WAL in it and pg_start_backup() will not switch the WAL segment if it is empty. These are both likely scenarios when setting up and/or testing pgBackRest. If the WAL segment is switched by pg_start_backup(), then check the archive for the segment that was detected prior to backup start. This should be common on normal running clusters with regular activity. Note that this might not be the segment immediately prior to the backup start segment if WAL volume is high. If pg_start_backup() did not switch the WAL then we can force a switch on PostgreSQL >= 9.3 by creating a restore point. In that case the WAL to check will be the backup start WAL. This is most likely to happen on idle systems, during testing, or immediately after a repo switch. An advantage of this approach other than earlier notification is that the backup directory will not be created so no resume will be attempted on the next backup. Note that some additional churn was created in backup.c because the load of archive.info needs to be done earlier.	2021-11-18 16:18:10 -05:00
David Steele	dea752477a	Remove obsolete statement about future multi-repository support.	2021-11-17 16:39:04 -05:00
Reid Thompson	1e77fc3d75	Include backup_label and tablespace_map file sizes in log output. In cases where they are returned by postgres, include backup_label and tablespace_map file sizes in the backup size value output in the log.	2021-11-16 10:21:32 -05:00
David Steele	df89eff429	Fix typos and improve documentation for the tablespace-map-all option.	2021-11-15 16:53:41 -05:00
David Steele	afe77e76e0	Update contributor for `6e635764`.	2021-11-10 07:31:02 -05:00
Reid Thompson	6e635764a6	Match backup log size with size reported by info command. Properly log the size of files copied during the backup, matching the backup size returned from the info command. In the reference issue, the incremental backup after switchover logs the size of all files evaluated rather than only the size of the files copied in the backup.	2021-11-09 13:24:56 -05:00
David Steele	038abaa71d	Display size option default and allowed values with appropriate units. Size option default and allowed values were displayed in bytes, which was confusing for the user. This also lays the groundwork for adding units to time options. Move option parsing functions into a common module so they can be used from the build module.	2021-11-03 15:23:08 -04:00
Reid Thompson	2a576477b3	Add --cmd option. Allows users to provide an executable to be used when pgbackrest generates command strings that expect to invoke pgbackrest. These generated commands are written to files by pgbackrest, e.g. recovery.conf.	2021-11-03 11:36:34 -04:00
David Steele	c5b5b58806	Simplify error handler. The error handler used a loop to process try, catch, and finally blocks. This worked fine but static analysis tools like Coverity did not understand that the finally block would always run and so there were false positives about double-free, unfreed resource, etc. This implementation removes the loop, which simplifies everything, and makes it clear that the finally block will always run. This cuts down on Coverity false positives. This implementation also catches lack of coverage on empty catch blocks so a few test fixes were committed separately in `d74fe7a`. A small refactor in backup.c is required because gcc 10.3.1 on Fedora 33 complains that the reason variable may be used uninitialized. It's not clear why this is the case, but reducing the scope of the TRY block fixes the issue.	2021-11-03 10:36:31 -04:00
David Steele	7f6c513be9	Add StringId as an option type. Rather the converting String to StringIds at runtime, store defaults in StringId format in parse.auto.c and convert user input to StringId during parsing.	2021-11-03 07:27:26 -04:00
David Steele	b13844086d	Use cfgOptionStrId() instead of cfgOptionStr() where appropriate. The compress-type, repo-type and log-level-* options have allow lists, which means it is more efficient to treat them as StringIds. For compress-type and log-level-* also update the functions that convert them to enums.	2021-11-01 17:35:19 -04:00
David Steele	bc352fa6a8	Simplify strIdFrom() functions. The strIdFrom() forced the caller to pick an encoding, which led to a number of TRY...CATCH blocks in the code. In practice the caller does not care which encoding is used as long as the string is valid for some encoding. Update the strIdFrom*() function to try all possible encodings and only throw an error when the string is not valid for any of them.	2021-11-01 10:08:56 -04:00
David Steele	904b897f5e	Begin v2.37 development.	2021-11-01 09:03:42 -04:00
David Steele	42fd6ce4e0	v2.36: Minor Bug Fixes and Improvements Bug Fixes: * Allow "global" as a stanza prefix. (Reviewed by Stefan Fercot. Reported by Younes Alhroub.) * Fix segfault on invalid GCS key file. (Reviewed by Stephen Frost. Reported by Henrik Feldt.) Improvements: * Allow link-map option to create new links. (Reviewed by Don Seiler, Stefan Fercot, Chris Bandy. Suggested by Don Seiler.) * Increase max index allowed for pg/repo options to 256. (Reviewed by Cynthia Shang.) * Add WebIdentity authentication for AWS S3. (Reviewed by James Callahan, Reid Thompson, Benjamin Blattberg, Andrew L'Ecuyer.) * Report backup file validation errors in backup.info. (Contributed by Stefan Fercot. Reviewed by David Steele.) * Add recovery start time to online backup restore log. (Reviewed by Tom Swartz, Stefan Fercot. Suggested by Tom Swartz.) * Report original error and retries on local job failure. (Reviewed by Stefan Fercot.) * Rename page checksum error to error list in info text output. (Reviewed by Stefan Fercot.) * Add hints to standby replay timeout message. (Reviewed by Cynthia Shang, Stefan Fercot. Suggested by Leigh Downs.)	2021-11-01 08:59:14 -04:00
David Steele	1336657326	Restore some linefeed rendering behavior from before `def7d513`. The new rendering behavior is correct in normal cases, but for the pre-rendered HTML blocks in the command and configuration references it causes a lot of churn. This would be OK if the new HTML was diff-able, but it is not. Go back to the old behavior of using br tags for this case to reduce churn until a more permanent solution is found.	2021-10-29 10:35:56 -04:00
David Steele	4f10441574	Add missing paragraph tags in coding standards.	2021-10-26 08:25:21 -04:00
David Steele	13d4559708	Check return value of getsockopt(). Checking the return value is not terribly important here, but if setsockopt() fails it is likely that bind() will fail as well. May as well get it over with and this makes Coverity happy.	2021-10-25 15:31:39 -04:00
Reid Thompson	1152f7a7d6	Fix mismatched parameters in tlsClientNew() call. `3879bc69` added this call and the parameters were not quite right but in way that the compiler decided they were OK. It was mostly working but TLS verification was disabled if caPath was NULL, which is not OK.	2021-10-25 12:56:33 -04:00
David Steele	3879bc69b8	Add WebIdentity authentication for AWS S3. This allows credentials to be automatically acquired in an EKS environment.	2021-10-22 18:31:55 -04:00
David Steele	51785739f4	Store config values as a union instead of a variant. The variants were needed to easily serialize configurations for the Perl code. Unions are more efficient and will allow us to add new types that are not supported by variants, e.g. StringId.	2021-10-22 18:02:20 -04:00
David Steele	2cea005f74	Fix segfault on invalid GCS key file.	2021-10-22 17:19:16 -04:00
David Steele	e443e3c6c0	Add br tags for HTML documentation rendering missed in `def7d513`.	2021-10-19 09:06:06 -04:00
David Steele	ccc255d3e0	Add TLS Server. The TLS server is an alternative to using SSH for protocol connections to remote hosts. This command is currently experimental and intended only for trial and testing. As such, the new commands and options will not show up in the command-line help unless directly requested.	2021-10-18 14:32:41 -04:00
David Steele	498902e885	Allow "global" as a stanza prefix. A stanza name like global_stanza was not allowed because the code was not selective enough about how a global section should be formatted. Update the config parser to correctly recognize global sections.	2021-10-07 12:18:24 -04:00
David Steele	68c5f3eaf1	Allow link-map option to create new links. Currently link-map only allows links that exist in the backup manifest to be remapped to a new destination. Allow link-map to create a new link as long as a valid path/file from the backup is referenced.	2021-10-05 17:59:05 -04:00
David Steele	6af827cbb1	Report original error and retries on local job failure. The local process will retry jobs (e.g. backup file) but after a certain number of failures gives up. Previously, the last error was reported but generally the first error is far more valuable. The last error is likely to be a cascade failure such as the protocol being out of sync. Report the first error (and stack trace) and append the retry errors to the first error without stack trace information.	2021-10-05 09:00:16 -04:00
Stefan Fercot	34f7873432	Report backup file validation errors in backup.info. Currently errors found during the backup are only available in text output when specifying --set. Add a flag to backup.info that is available in both the text and json output when --set is not specified. This at least provides the basic info that an error was found in the cluster during the backup, though details are still only available as described above.	2021-10-04 13:45:53 -04:00
David Steele	71047a9d6d	Use strncpy() to limit characters copied to optionName. Valgrind complained about uninitialized values on arm64 when comparing the reset prefix, probably because "reset" ended up being larger than the option name: Conditional jump or move depends on uninitialised value(s) at cfgParseOption (parse.c:568). Coverity complained because it could not verify the size of the string to be copied into optionName, probably because it does not understand the purpose of strSize(): You might overrun the 65-character fixed-size string "optionName" by copying the return value of "strZ" without checking the length. Use strncpy() even though we have already checked the size and make sure the string is terminated. Keep the size check because searching for truncated option names is not a good idea. This is not a production bug since the code has not been released yet.	2021-10-02 16:17:33 -04:00
David Steele	9e79f0e64b	Add recovery start time to online backup restore log. This helps give an idea of how much recovery needs to be done to reach the end of the WAL stream and is easier to read than the backup label.	2021-09-29 10:31:51 -04:00
David Steele	9346895f5b	Rename page checksum error to error list in info text output. "error list" makes it clearer that other errors may be reported. For example, if checksum-page is true in the manifest but no checksum-page-error list is provided then the error is in alignment, i.e. the file size is not a multiple of the page size, with allowances made for a valid-looking partial page at the end of the file. It is still not possible to differentiate between alignment and page checksum errors in the output but this will be addressed in a future commit.	2021-09-29 09:58:47 -04:00
David Steele	b7ef12a76f	Add hints to standby replay timeout message.	2021-09-28 15:55:13 -04:00
David Steele	def7d513cd	Eliminate linefeed formatting from documentation. Linefeeds were originally used in the place of <p> tags to denote a paragraph. While much of the linefeed usage has been replaced over time, there were many places where it was still being used, especially in reference.xml. This made it difficult to get consistent formatting across different output types. In particular there were formatting issues in the command-line help because it is harder to audit than HTML or PDF. Replace linefeed formatting with proper <p> tags to make formatting more consistent. Remove double spaces in all text where <p> tags were added since it does not add churn. Update all <ul>/<ol>/<li> tags to the more general <list>/<list-item> tags. Add a few missing periods.	2021-09-08 17:35:45 -04:00
David Steele	02b06aa495	Increase max index allowed for pg/repo options to 256. The prior limitations were based on using getopt_long() to parse command-line options, which required a static list of allowed options. Setting index max too high bloated the binary unacceptably. `45a4e80` replaced the functionality of getopt_long() but the static list remained. Improve cfgParseOption() to use available option data and remove the need for a static list. This also allows the option deprecations to be represented more compactly. Index max is still capped at 256 because a large enough index could cause parseOptionIdxValue() to run out of memory since it allocates a static list based on the highest index found. If that function were improved with a map of found index values then index max could be set to UINT64_MAX. Note that deprecations no longer set an index max or define whether reset is valid. These were space-saving measures which are no longer required. This means that indexed deprecated options will also be valid up to 256 and always allow reset, but it doesn't seem worth additional code to limit this behavior. cfgParseOptionId() is no longer needed because calling cfgParseOption() with .ignoreMissingIndex = true duplicates the functionality of cfgParseOptionId(). This leads to some simplification in the help code.	2021-08-31 12:09:50 -04:00
David Steele	aee0e7bac7	Begin v2.36 development.	2021-08-23 07:03:40 -04:00
David Steele	3787cf7803	v2.35: Binary Protocol IMPORTANT NOTE: The log level for copied files in the backup/restore commands has been changed to detail. This makes the info log level less noisy but if these messages are required then set the log level for the backup/restore commands to detail. Bug Fixes: * Detect errors in S3 multi-part upload finalize. (Reviewed by Cynthia Shang, Marco Montagna. Reported by Marco Montagna, Lev Kokotov, Anderson A. Mallmann.) * Fix detection of circular symlinks. (Reviewed by Stefan Fercot. Reported by Rohit Raveendran.) * Only pass selected repo options to the remote. (Reviewed by David Christensen, Cynthia Shang. Reported by Greg Sabino Mullane, David Christensen.) Improvements: * Binary protocol. (Reviewed by Cynthia Shang.) * Automatically create data directory on restore. (Contributed by Stefan Fercot. Reviewed by David Steele. Suggested by Chris Bandy.) * Allow restore --type=lsn. (Contributed by Stefan Fercot. Reviewed by Cynthia Shang. Suggested by James Coleman.) * Change level of backup/restore copied file logging to detail. (Reviewed by Stefan Fercot. Suggested by Jens Wilke.) * Loop while waiting for checkpoint LSN to reach replay LSN. (Contributed by Stefan Fercot. Reviewed by David Steele. Suggested by Fatih Mencutekin.) * Log backup file total and restore size/file total. (Reviewed by Cynthia Shang.) Documentation Bug Fixes: * Fix incorrect host names in user guide. (Reviewed by Stefan Fercot. Reported by Greg Sabino Mullane.) Documentation Improvements: * Update contributing documentation and add pull request template. (Contributed by Cynthia Shang. Reviewed by David Steele.) * Rearrange backup documentation in user guide. (Reviewed by Cynthia Shang.) * Clarify restore --type behavior in command reference. (Contributed by Cynthia Shang. Reviewed by David Steele.) * Fix documentation and comment typos. (Contributed by Eric Radman. Reviewed by David Steele.) Test Suite Improvements: * Add check for test path inside repo path. (Reviewed by Greg Sabino Mullane. Suggested by Greg Sabino Mullane.) * Add CodeQL static code analysis. (Reviewed by Cynthia Shang.) * Update tests to use standard patterns. (Contributed by Cynthia Shang. Reviewed by David Steele.)	2021-08-23 06:52:51 -04:00
David Steele	4fb6384f10	Fix more memory leaks introduced by the binary protocol in `6a1c0337`. Either of these temp mem context blocks fixes the issue of command packs not being freed, but it seems like a good idea to have both in case the code changes.	2021-08-18 08:18:11 -04:00
Cynthia Shang	eca2fc6958	Update config/parse test to use standard patterns.	2021-08-12 12:38:07 -04:00
Cynthia Shang	e17865a03a	Update protocol/protocol test to use standard patterns.	2021-08-12 11:57:17 -04:00
David Steele	a0bdfa436c	Log backup file total and restore size/file total. The backup size was a bit off because it did not include any files (e.g. backup_label, WAL files) that were added to the manifest after the main copy. To fix this move the log message to the very end of the backup. Add size/file total log message to restore since it did not exist before.	2021-08-11 13:39:36 -04:00
David Steele	6ab18dc0fa	Rearrange backup documentation in user guide. Remove the "Automatic Stop Option" section since it only applies to PostgreSQL <= 9.6, which will soon be EOL. Since we no longer build the user guide for PostgreSQL < 10 this section was no longer being tested. The stop-auto option is still documented in the reference. Move the "Fast Start Option" to "Quick Start - Perform Backup". This is a commonly-used option so it makes sense to mention it earlier. This also makes the backups run more quickly. In the worst case, backups in "Quick Start - Perform Backup" could take minutes to start Move the "Archive Timeout" section to "Quick Start - Perform Backup" since it is the last section in "Backup".	2021-08-11 12:59:25 -04:00
David Steele	f716cb6f4f	Fix use after free introduced by the binary protocol in `6a1c0337`. The user and group were stored in a temp reset mem context so they could get freed if there were enough files to trigger the reset in storageRemoteInfoList(). Allocate user and group in a mem context provided by the caller to prevent them being freed prematurely.	2021-08-10 14:22:38 -04:00
Cynthia Shang	71b654fc29	Fix links and update child process example. Removed colon from example titles to fix links, fixed test.yml link, and updated the example for the parent/child test process to use the latest macros instead of sleep().	2021-08-09 16:56:06 -04:00
Cynthia Shang	f653b59664	Update db/db test to use standard patterns.	2021-08-09 16:35:48 -04:00

1 2 3 4 5 ...

1484 Commits