A file entry in a zip whose name refers to a directory, such as ".",
"/", "" or "sub/.", was only skipped when it named the root of an
archive which was itself the root of the remote. When the archive was
found by listing its parent directory the entry appeared as a file
with the same name as the archive alongside the directory for it, and
copying the archive tried to write both. When the entry named a
subdirectory it appeared as a file alongside that directory, and with
that subdirectory mounted as the archive root the entry was taken to
be the single file the root points at, hiding every real entry.
Skip any file entry whose last path component is "", "." or "..",
checked on the raw name before it is cleaned or joined on the prefix.
(cherry picked from commit da352a2a1b)
The zip archiver handed out its cached directory tree directly. Any
caller which filters a listing in place (as the core listing code
does) altered the cache, so later listings of the same directory could
be corrupted.
Return a copy of the cached listing instead.
(cherry picked from commit 68eab60564)
When an upload failed with a 500 error the upload was retried with a
new upload link but the same input stream. The stream had already been
consumed by the first attempt so the retry uploaded an empty file.
This fixes it by returning a RetryError instead so the caller retries
the upload with a fresh stream, which will fetch a new upload link.
(cherry picked from commit 6cdd0ea761)
When an upload failed with a retryable error the pacer retried the
whole upload with the same input stream. The stream had already been
consumed by the first attempt so the retry uploaded an empty file.
This fixes it by using CallNoRetry for the upload, as the other
backends do, so retryable errors are returned wrapped in a RetryError
for the caller to retry the upload with a fresh stream.
It also makes 5xx errors from the upload storage servers retryable.
These come back from the SDK as a different error type to API errors
so were not being retried at all.
(cherry picked from commit fa43f10af2)
When an upload failed with a retryable error the pacer retried the
whole PUT with the same input stream. The stream had already been
consumed by the first attempt so the retry uploaded an empty file.
This fixes it by using CallNoRetry for the upload, as the other
backends do, so retryable errors are returned wrapped in a RetryError
for the caller to retry the upload with a fresh stream.
(cherry picked from commit f4cd80a535)
The headers set with --http-headers are documented for passing
credentials such as Authorization or Cookie. The backend used the
default net/http redirect policy which copies all but a handful of
well known headers to any redirect target, so a redirect from the
configured server to another host would send those credentials to
that host, and a redirect from https to http would send them in
plaintext.
When headers are configured this installs a CheckRedirect function
which:
- removes the configured headers from every hop once the redirect
chain has left the originally requested host
- refuses a redirect from https to http with an error
(cherry picked from commit 0adc0082dff7ef6ed1597418f85ac691478a6fa5)
Whether an archive entry name can escape the archive's namespace was
left entirely to each archiver. Enforce it in the archive backend too.
List only passes on direct children of the directory listed and
NewObject only returns the object asked for, so a future archiver
which forgets to validate names cannot expose a traversal to fs/sync
and fs/operations.
(cherry picked from commit a6a7d95e081b91231b49c4004e8a2bff16da4cd8)
The path inside the archive was compared against the cleaned entry
names without being cleaned itself, so `archive.zip/sub/./dir` or
`archive.zip/sub//dir` failed to list even though `archive.zip/sub/dir`
worked.
(cherry picked from commit 34636f34079585f3a8f883038483ece58d39b084)
A zip containing a file entry whose name refers to the archive's own
root (".", "/" or "") was presented as a single file called "." and
all its other entries disappeared. A file at the root can only be the
archive member the backend was pointed at, so with no root such an
entry is skipped like any other unsafe name.
(cherry picked from commit c9fac578aae7707faf340b58cf45465a5f698cbd)
Entry names read from a squashfs directory are not sanitized by
go-diskfs. The squashfs backend joined each leaf name onto its
directory to form the object's remote, so a crafted image could escape
its directory.
Use sanitize.Leaf to skip unsafe entries in List. A "\" is an
ordinary character in a file name on the systems squashfs images are
made on and in an rclone remote path, so it is deliberately not
rejected; making it safe for the destination is the destination
backend's job.
Skipped entries are logged at DEBUG with a single NOTICE count per
listing so a crafted image under a mount cannot flood the log.
(cherry picked from commit 63385c66175141b63bb53b3e42b50908649f8437)
When a zip archive was mounted at a subdirectory root, readZip used a bare
strings.HasPrefix to decide which entries fell inside the root. This
matched on a raw string prefix rather than a path boundary, so mounting
root "foo" also exposed sibling entries such as "foobar/..." with their
names left uncorrected.
Require a path boundary when filtering by root.
(cherry picked from commit b444227264cee2a9307f03ba90c2a411e7eeb693)
The zip backend mounts a zip file as a browsable Fs. Go's archive/zip
does not sanitize entry names, and readZip applied path.Clean but did
not reject a cleaned name that still pointed outside the archive. A
crafted zip could make rclone copy/sync attempt writes outside the
intended destination.
Sanitize entry names with sanitize.Path - the same check used by
rclone archive extract - skipping any entry with a ".." path
component, whether separated by "/" or "\". A backslash is otherwise
kept as an ordinary character in the name, as archive extract does. It
is up to the destination backend to make names safe for its storage.
Skipped entries are logged as a single count per archive so a crafted
archive with many escaping entries cannot flood the log.
(cherry picked from commit 224fe97ac13105094651264a9e16ee3cf41ccf77)
With --links/-l, a symlink is served as a .rclonelink object whose
content is the target path. A Range request with a start offset beyond
the target length (e.g. "Range: bytes=99999999999-") reached
openTranslatedLink and sliced the target string at that offset, panicking
with "slice bounds out of range".
Clamp the offset to the target length so an out-of-range start reads
empty, matching how a real file read past EOF behaves.
(cherry picked from commit 3fa32192c20f0b4cfd4f4b07636dc3445026041a)
The birth-time (btime) write in writeMetadataToFile followed symlinks for
any object that was not a translated link, so under -l/--links a symlink
planted by an untrusted source at the destination path could redirect the
btime write to a target outside the backup destination on OSes where
birth time is settable (Windows).
Use the NOFOLLOW birth-time write whenever translating symlinks, not only
for translated links. It is a no-op on a real file or directory and stops
a planted symlink from being followed out of the destination.
(cherry picked from commit bd86336faac7cfc4e9057add38ec5fd12d762d02)
With -l/--links the local backend faithfully recreates a source ".rclonelink" as
a real symlink at the destination. Directory metadata (chmod/chown/chtimes),
however, was applied with the raw following syscalls
os.Chmod/os.Chown/os.Chtimes rather than through the os.Root sandbox used for
content writes. A Directory is never a translatedLink, so when the destination
path already existed as a symlink planted by an untrusted source, the metadata
was applied through it to a target outside the backup destination.
Route directory metadata through os.Root when translating symlinks, so a planted
symlink can no longer redirect chmod/chown/chtimes out of the destination, while
legitimate in-tree directories are unaffected.
(cherry picked from commit b764ccbd23582f29c611862f11f86dde3544035a)
Each upload chunk is buffered in a pool.RW from the global memory pool
but was never closed, so its pages were never returned to the pool.
Close the buffer after each chunk is uploaded and on the read error
path.
A chunk that failed with a retryable error was also retried without
rewinding the buffer, so the retry sent an empty body with the original
Content-Length and Content-Range and failed.
Seek the chunk back to the start inside the pacer closure so each
attempt re-sends it in full.
The FsPutRetry integration test covers the retry of a failed upload
request and checks the buffers are returned to the pool.
(cherry picked from commit 2f0228029e)
When both /me/drives and /me/drive fail during config (for example an
account-level 403 serviceReadOnly "Database Is Read Only"), send the
config state machine to the existing manual drive ID entry state
instead of dead-ending at choose_type with the raw error. The drive
itself remains usable when only the enumeration API is blocked.
Fixes#9794
(cherry picked from commit 03fe2ef794)
Directory names which look like they have a --b2-versions version
string are now encrypted in full, so directories created by older
rclone (which left the version string in plain text) no longer
decrypt and vanished silently from listings.
DecryptDirName now falls back to the old form for such names so the
directory is listed, and logs the name it needs to be renamed to on
the underlying remote to make it accessible again. Document this in
the crypt docs.
(cherry picked from commit 1583cce1e2)
The refresh endpoint returns a rotated token with a fresh expiry on
every successful call, but getUserInfo discarded it, so routine use
never extended the stored token's life. Once the stored token aged
out, accounts with 2FA enabled could not recover non-interactively
and required a manual reconnect.
Carry the rotated token out of getUserInfo and persist it in NewFs
via the same jwtToOAuth2Token + oauthutil.PutToken path that
refreshJWTToken uses, keeping f.cfg.Token in sync (same pattern as
refreshOrReLogin).
Fixes#9584
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 66761670da)
The --b2-versions support added in 3fe2aaf96 strips a version string
from the last segment of a path before encrypting it, so that the
plain text version suffixes which the underlying backend appends to
encrypted file leaf names can be handled. EncryptDirName and
DecryptDirName share that code, so the last segment of a *directory*
name was version stripped too. Only file leaf names are ever given a
version string by the backend - a directory gets a
version-string-like name from the user, and such a name is encrypted
verbatim when it appears as the parent of a file name, so the same
directory ended up with two different encryptions.
Before this change, with a directory whose name matches rclone's
version format, eg dir-v2001-02-03-040506-123:
rclone copy file.txt crypt:dir-v2001-02-03-040506-123/
rclone ls crypt:dir-v2001-02-03-040506-123
# => "directory not found" - the file is invisible to listings
rclone mkdir crypt:dir-v2001-02-03-040506-123
# => creates a second directory with the same decrypted name
After this change EncryptDirName and DecryptDirName encrypt directory
names verbatim, so a directory encrypts the same way whether it is
named on its own or as the parent of a file. Version strings are only
added to file names by the underlying backend, so --b2-versions is
unaffected and the existing version tests are untouched.
A directory which was created by the old EncryptDirName will no longer
decrypt and will be reported as undecryptable in listings. Such
directories were already unusable - anything copied into one was
written to a different encrypted directory - so nothing which worked
before is broken by this.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 67b184d6e7)
With no_head_object set, NewObject does not read any metadata, so the
destination object returned from a server side copy had a size of 0.
The size check in operations.Copy then failed with "corrupted on
transfer: sizes differ N vs 0" and deleted the newly copied object.
This also broke Move and hence renames through rclone mount.
Populate the destination object's size and MD5 from the source object
when no_head_object is set, as a server side copy produces an object
with identical content.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 6df7b8aba1)
Dropbox is case insensitive and the path_display it returns in
change notifications may not match the case of the configured root.
Before this change the root was trimmed with a case sensitive prefix
match, so when the cases differed the full path was passed to the
ChangeNotify callback and the notification was ignored.
This trims the root case insensitively while preserving the display
case of the remaining path.
(cherry picked from commit 4af64270cc)
A successful UploadPart whose response carries no ETag header made
WriteChunk panic dereferencing uout.ETag in a debug log line. The part
ETag is required by CompleteMultipartUpload, so an ETag-less 200 is
unusable: return a retryable error from inside the pacer callback so
the chunk is retried instead of crashing the transfer or completing
the upload with a broken part list.
Fixes#9822
Co-authored-by: Shurong Cao <170531907+CAOShurong@users.noreply.github.com>
(cherry picked from commit 660144d311)
Nextcloud only stores a checksum which is supplied in the OC-Checksum
header of an upload, and discards it again when the modification time
is set with PROPPATCH. Re-sending the checksum in the PROPPATCH (as is
done for ownCloud) is rejected by Nextcloud with 403 Forbidden which
made the whole PROPPATCH fail, so SetModTime returned an error on any
object which had a hash. Uploads from sources without hashes, eg
streamed uploads with `rclone rcat`, were stored with no hash at all.
Use the Nextcloud PATCH extension with the X-Recalculate-Hash header
to have the server calculate and store the SHA1 of an object after a
streamed upload and after setting the modification time. This gives
a server side hash of the stored data which also lets rclone verify
streamed uploads.
(cherry picked from commit f7c510af49)
If the source supplied fewer bytes than its declared size, the single
part upload path accepted the short body and stored a truncated file
recorded with the declared size, reporting a successful upload. The
multipart path already checks for this.
Count the bytes actually read and fail the upload if they do not match
the declared size, which cancels the partially created file.
This was found by the FsPutShortEOF integration test.
(cherry picked from commit 8e744de5e6)
We accidentally merged this commit with non compiling tests.
bee45bccfd azureblob: fix spurious vfs cache corruption errors during chunked reads #9782
(cherry picked from commit d3a71eea36)
listSharedFolders already decoded shared-folder names with
f.opt.Enc.ToStandardName, but listReceivedFiles stored the raw name
returned by the Dropbox API unchanged. Names that require encoding
(e.g. a trailing space, which Dropbox itself rejects, so rclone
stores it as "name␠" via EncodeRightSpace) were therefore shown under
their raw, encoded form for received files instead of being decoded
back to the standard name, and findSharedFile could not resolve such
a file by its standard name.
Apply the same ToStandardName conversion listSharedFolders uses.
(cherry picked from commit f3a7aaf635)
On a ranged download the metadata decoder stored the response's
Content-Length (the length of the range, not the blob) in the object's
size and only corrected it from the Content-Range total afterwards.
Object.Size() is read concurrently by the VFS cache and chunked reader
while a download is in progress, so with --vfs-read-chunk-size a reader
could observe the chunk length (e.g. 67108864 for 64M chunks) as the
object size. The VFS cache then logged
vfs cache: cached file (N) is unexpectedly larger than the remote
object (67108864). The cached file is likely corrupted after an
unclean shutdown; recovering ...
and truncated the read request against the bogus size, breaking
sequential reads of large blobs with --vfs-cache-mode full.
This applies the Content-Range correction before the size is stored so
the range length is never published as the object size.
(cherry picked from commit bee45bccfd)
The multipart upload copied the source into the request buffer without
checking how many bytes it had read, so a source that supplied fewer
bytes than its declared size was accepted by the server and reported as
a success with a truncated file stored.
Count the bytes actually read and fail the upload if they do not match
the declared size.
Signed-off-by: Rohit Behera <126186063+r0h1tb@users.noreply.github.com>
(cherry picked from commit 83b143103c)
The single-shot upload path sent the source straight to Box as a multipart
body with no Content-Length, so a source that supplied fewer bytes than its
declared size produced a short request that Box accepted and stored, and the
upload was reported as a success.
Count the bytes actually read and fail the upload if they do not match the
declared size. The multipart path already reads each chunk with io.ReadFull
and so already fails in this case.
(cherry picked from commit 64ab1ac322)
The concurrent walker created by walk() only stopped when the callback
returned an error or the whole tree had been listed. Cancelling the
context (for example via the rc job/stop endpoint for an async
operations/size or recursive operations/list call) was therefore
ignored: the checkers kept pulling list jobs from the channel and kept
listing the entire tree, burning CPU and making job cancellation
useless for every backend without a native ListR implementation.
Make every checker select on ctx.Done() so a cancelled walk shuts down
promptly through the existing quit/drain path and reports the context
error. Also check the context between directory read chunks in the
local backend so a single huge directory does not block cancellation.
(cherry picked from commit 5eb5c01e36)
Update already wrapped the source in a counting reader but never looked at the
count, so a source that supplied fewer bytes than its declared size was
uploaded as a chunked request, accepted by the server and reported as a
success with a truncated file stored.
Compare the bytes actually read against the declared size.
(cherry picked from commit 1128693468)
Before this change, when no_data_encryption was set, uploads from
local disk advertised the hash of the encrypted data even though the
data was uploaded unencrypted.
On backends which check upload hashes (eg b2) this made uploads of
small files fail with errors like "Checksum did not match data
received", and made chunked uploads store an incorrect hash so the
files failed their checksum on download with "corrupted on transfer:
SHA1 hashes differ".
See: https://forum.rclone.org/t/sha1-mismatches-on-b2-with-no-data-encryption-true/54121
(cherry picked from commit 5a0b7d6746)
Writing any file into a third-party app container - the Obsidian, Pages or
Shortcuts folders that iCloud Drive shows alongside your own - failed with
HTTP error 412 (412 Precondition Failed) returned body:
"{ ... \"error_code\" : \"VALIDATING_REFERENCE_ERROR\", \"reason\" :
\"Request has out of order children to be chained but the parents were
missing\" }"
Reading from those paths worked, and so did creating directories in them, so
the failure looked like a missing parent when the parent was plainly there.
Items in an app container live in a different zone from ordinary iCloud Drive
folders: a folder under Documents has a drive ID like
FOLDER::com.apple.CloudDocs::<uuid>, while the Obsidian container has
FOLDER::iCloud.md.obsidian::documents#o2v. DownloadFile already accounts for
this - it deconstructs the item's own ID and addresses the zone it finds - but
CreateUpload and UpdateFile hardcoded defaultZone, and UpdateFile built the
resulting Drivewsid with a hardcoded com.apple.CloudDocs as well.
So rclone asked Apple to chain the new document to a parent in
com.apple.CloudDocs while the parent lived in iCloud.md.obsidian. The parent
really was missing from the zone being addressed, which is what the error said.
Take the zone from the parent's drive ID instead, the same way the download
path does, and build the new item's ID with ConstructDriveID. Uploads outside
an app container are unaffected: their parents are in com.apple.CloudDocs, so
the derived zone is the value that was previously hardcoded.
Verified against a real remote: files now upload into an Obsidian vault inside
the container and read back correctly with an unpatched binary afterwards.
(cherry picked from commit aba403fa79)
If the source supplied fewer bytes than its declared size, the upload
request failed but a retry could report success even though the stored
file was truncated, because the retry re-sent an already exhausted
reader.
Count the bytes actually read from the source and if they do not match
the declared size return an error.
This was found by the new FsPutShortEOF integration test.
(cherry picked from commit 7357fb82a9)
If the source supplied fewer bytes than its declared size, the upload
request failed but a retry could report success even though the stored
file was truncated, because the retry re-sent an already exhausted
reader.
Count the bytes actually read from the source and if they do not match
the declared size, remove the partially uploaded file and return an
error.
This was found by the new FsPutShortEOF integration test.
(cherry picked from commit 5d057549b9)
If the source supplied fewer bytes than its declared size, the
multipart upload was completed anyway, storing a truncated file and
reporting a successful upload.
Check the number of bytes read from the source against the declared
size before finalising and abort the upload with an error if they do
not match.
This was found by the new FsPutShortEOF integration test.
(cherry picked from commit a2baa978db)
If the source supplied fewer bytes than its declared size, the
truncated file was stored and the upload reported success with the
object claiming the declared size.
Count the bytes actually read from the source and if they do not match
the declared size, remove the truncated file and return an error.
This was found by the new TestRcatSizeShortEOF integration test.
(cherry picked from commit e1bf9405e2)
The file is created at the declared size and the data then written
with ranged writes, so if the source supplied fewer bytes than
declared, the remainder of the file was left as zeroes and the upload
reported success.
Count the bytes actually read from the source and if they do not match
the declared size, delete the partially uploaded file (if newly
created) and return an error.
This was found by the new FsPutShortEOF and TestRcatSizeShortEOF
integration tests.
(cherry picked from commit 884b28c203)
When using Microsoft Entra ID credentials, Azure Blob server-side copy
uses a user delegation SAS URL for the private copy source.
The SAS start time was set to the current local time. Azure Storage
validates the copy source from the service side, and small clock
differences can make that SAS appear not yet valid. The service then
returns 403 CannotVerifyCopySource with AuthenticationFailed.
Start the copy-source SAS 15 minutes in the past, matching Microsoft SAS
guidance for clock skew.
(cherry picked from commit e4c7aca6bd)
The precision field in the backend overview YAML files can hold
fs.ModTimeNotSupported (100 years in nanoseconds) which overflows int
on 32 bit platforms, making the YAML for those backends fail to parse
and causing rclone to log 18 internal errors on every invocation.
Use int64 for the precision field and add a test that parses every
embedded backend YAML file so this is caught on 32 bit test runs.
(cherry picked from commit 5629f2668c)
When an upload failed part way through with a retryable error (eg a
502 from the block storage servers) the pacer retried the whole upload
call with the same input stream. The stream had already been partially
consumed, so the retry re-created the upload draft and committed just
the remainder of the stream as a complete file, silently truncating
it. With restic over serve restic this corrupted the repository as the
truncated pack was reported as successfully uploaded.
This fixes it by using CallNoRetry for the upload, as the other
backends do, so retryable errors are returned wrapped in a RetryError
for the caller to retry the upload with a fresh stream.
(cherry picked from commit a06df7a2de)
When the source supplied fewer bytes than its declared size, the
compressed data file was stored under a name containing the declared
size while the metadata recorded the actual number of bytes read.
NewObject looks the data file up by the size in the metadata, so the
resulting object could never be read again, and the upload reported
success.
Check that the number of bytes read matches the declared size after
uploading the data and before writing the metadata, and remove the
data file and return an error if it does not.
This was found by the new FsPutShortEOF integration test.
(cherry picked from commit 18fa445ffc)
The append loop retries everything once the upload session has
started, so a cancelled context error was retried through all the low
level retries with exponential backoff before the upload gave up.
(cherry picked from commit e0701daea0)
A source which returned EOF before supplying as many bytes as it
declared would either commit a truncated file (if the shortfall was
within the final chunk) or loop forever appending empty chunks to the
upload session. Return an error wrapping io.ErrUnexpectedEOF instead.
Note that all dropbox uploads use the chunked upload path with the
default batch_mode of sync, so this affected uploads of every size.
(cherry picked from commit bff17664ad)
Object.Update held its connection until the deferred putConnection ran at
function exit, so the SetModTime it does at the end of every upload had to take
a second connection from the pool, dialling a whole new SMB session when the
pool was empty. With N transfers in flight the pool grew to roughly 2N sessions
for no reason.
Return the connection as soon as the file is closed. At that point the upload
has succeeded and remove() can no longer be reached, so nothing else needs it,
and SetModTime picks the same connection straight back out of the pool.
putConnection nils the pointer, so the deferred putConnection becomes a no-op
and the connection is not returned twice.
(cherry picked from commit ea9a64c751)