Skip to content

ScoutFS Backend

Ben McClelland edited this page Sep 18, 2026 · 23 revisions

The ScoutFS backend maps S3 buckets and objects onto a ScoutFS filesystem. It builds on the POSIX backend while adding:

  • Multipart completion through ScoutFS extent moves, avoiding a full read-and-write copy when the parts meet the filesystem requirements.
  • ScoutFS project IDs for quota accounting.
  • Optional S3 Glacier-compatible behavior for data managed by ScoutAM.
  • The ScoutAM noarchive flag on temporary multipart parts.

Use the scoutfs backend for a ScoutFS mount even when Glacier integration is disabled.

Platform and filesystem requirements

The backend is available only in Linux AMD64 builds. On other platforms, startup returns scoutfs only available on linux.

The gateway root must reside on a mounted ScoutFS filesystem. The gateway uses ScoutFS ioctls, system attributes, user xattrs, and normal POSIX file operations. The process needs the filesystem permissions and ScoutFS privileges required for the enabled features.

Metadata is always stored as xattrs. Unlike the POSIX backend, ScoutFS does not expose --sidecar or --nometa.

Command and gateway root

versitygw [global options] scoutfs [command options] <gateway-root>

The first positional argument is required. It must already exist, be a directory, and be accessible to the gateway process:

versitygw --access "$ROOT_ACCESS_KEY_ID" --secret "$ROOT_SECRET_ACCESS_KEY" \
  scoutfs --projectid /mnt/scoutfs/gateway

Every top-level directory below the root is a bucket. The backend changes the process working directory to the resolved gateway root during startup, so use absolute paths for --versioning-dir, access logs, certificates, and other file-based options.

Option summary

Option Default Purpose
--glacier, -g Disabled Report offline data as GLACIER and accept restore requests
--chuid Disabled Set ownership of new paths from the IAM account UID
--chgid Disabled Set ownership of new paths from the IAM account GID
--projectid Disabled Set the ScoutFS project ID from the IAM account ProjectID
--bucketlinks Disabled Include top-level directory symlinks in bucket discovery
--versioning-dir Disabled Store non-current object versions in a shadow namespace
--dir-perms 0755 Permission bits requested for newly created directories
--file-perms 0644 Permission bits applied to new object and multipart-part files
--disable-noarchive Disabled Stop marking temporary multipart parts as noarchive
--object-lock-mode none Select local or disabled conditional-publish locking
--concurrency 5000 Limit concurrent filesystem actions
--default-etag Empty Supply an ETag for files without stored ETag metadata
--data-integrity-etag Disabled Use checksum-derived instead of MD5-compatible ETags

Glacier emulation

--glacier, -g
VGW_SCOUTFS_GLACIER=true

--glacier inspects ScoutFS offline extents and translates their state into S3 storage-class, restore-status, and InvalidObjectState behavior. A ScoutAM policy must archive and stage the data; VersityGW only exposes the S3 workflow and sets the ScoutFS stage-request flag.

See Glacier mode for request behavior.

UID and GID ownership

--chuid
--chgid
VGW_CHOWN_UID=true
VGW_CHOWN_GID=true

These options set ownership on paths created by the gateway from the authenticated account's UID and GID. They affect new buckets, object parent directories, objects, and multipart parts, but do not recursively change existing data.

The process normally needs root or CAP_CHOWN to assign IDs different from its own. Startup warns when either option is set by a non-root process, and writes for accounts with different IDs fail unless the process can chown. The flags are independent; an ID not selected by a flag remains the process's effective UID or GID when a new path is chowned.

The standalone IAM service has no per-user POSIX identity. Every account receives the gateway-wide --iam-standalone-default-uid/--iam-standalone-default-gid values, so --chuid and --chgid cannot provide distinct ownership for those users.

ScoutFS project IDs

--projectid
VGW_SET_PROJECT_ID=true

This option applies a positive IAM account ProjectID to newly created buckets, object files, parent directories, multipart parts, and completed multipart objects. ScoutFS can then account and enforce storage by project. Zero and negative project IDs are ignored.

Project IDs require ScoutFS filesystem format version 2 or newer. At startup, the gateway reads /sys/fs/scoutfs/<filesystem-id>/format_version. If the format is older or cannot be confirmed, it prints a warning and disables project-ID assignment while continuing to serve requests.

The standalone IAM backend supplies one process-wide --iam-standalone-default-project-id for every account rather than distinct per-user project IDs.

Bucket symlinks

--bucketlinks
VGW_BUCKET_LINKS=true

--bucketlinks includes top-level symlinks to directories in ListBuckets. Object operations below such a name follow the symlink target, which can expose data outside the gateway root. This option controls discovery, not access control: disabling it does not guarantee that a direct request naming an existing symlink is blocked. Use only trusted links and protect targets with filesystem permissions.

Object listings do not recursively traverse object-level symlinks. A direct GET follows an object symlink, while DELETE removes the symlink.

Versioning storage

--versioning-dir <path>
VGW_VERSIONING_DIR=<path>

Versioning is experimental and requires an existing directory outside the gateway root. Startup rejects a missing path, a non-directory, or a directory contained by the gateway root. Current object versions remain in the primary namespace; non-current versions and delete markers use the shadow namespace.

Configuring the directory enables versioning capability, but clients must still set each bucket to Enabled or Suspended. Use storage with suitable capacity and durability. See POSIX Object Versioning.

Creation permissions

--dir-perms <octal-mode>
--file-perms <octal-mode>
VGW_DIR_PERMS=<octal-mode>
VGW_FILE_PERMS=<octal-mode>

--dir-perms supplies the requested mode for new buckets, object parent directories, multipart directories, versioning paths, and internal directories. Its default is 0755; the process umask can remove bits during directory creation.

--file-perms applies explicitly to new object and multipart-part files, so the process umask does not reduce it. The default is 0644, and values outside 0000 through 0777 are rejected. Neither option changes existing paths. Extremely restrictive modes can prevent later gateway operations.

Multipart noarchive flag

--disable-noarchive
VGW_DISABLE_NOARCHIVE=true

By default, each uploaded multipart part receives the ScoutAM noarchive flag. Parts are temporary and disappear after completion or abort, so archiving them wastes resources and can interfere with extent-move completion. If the flag cannot be set, UploadPart fails rather than silently leaving an archive-eligible part.

Use --disable-noarchive only when ScoutAM is absent or its policy requires parts to remain archive-eligible. The option does not change archive policy for completed objects.

Conditional-publish locking

--object-lock-mode <local|none>
VGW_OBJECT_LOCK_MODE=<local|none>

This option serializes the final condition check and publication for PutObject and CompleteMultipartUpload requests carrying If-Match or If-None-Match. It is unrelated to S3 Object Lock retention and legal hold.

Mode Behavior
none Default. Conditional object writes return NotImplemented; unconditional writes continue to work
local Serializes competing publishes within one gateway process

ScoutFS does not offer the POSIX backend's flock or fcntl shared modes. local therefore cannot make conditional writes atomic across multiple gateway processes sharing a filesystem. Any other mode is rejected at startup.

Filesystem concurrency

--concurrency <count>
VGW_POSIX_CONCURRENCY=<count>

The value must be positive and defaults to 5000. It bounds concurrently running ScoutFS/POSIX backend actions and provides backpressure when filesystem-heavy syscalls block. Requests wait when all slots are occupied, so sustained saturation increases latency and can cause client or proxy timeouts.

Tune this with the mounted filesystem, thread growth, file-descriptor limits, and request latency rather than CPU count alone.

Default ETag

--default-etag <value>
VGW_DEFAULT_ETAG=<value>

Files created directly in ScoutFS do not have VersityGW's stored ETag attribute. This option returns one constant fallback for those files during GET, HEAD, listing, copy, and precondition processing. When unset, their ETag is empty.

The value is not derived from file contents and can make cache validation or ETag preconditions imprecise. It is returned verbatim, so include quotation marks if clients require them; for example:

--default-etag '"external"'

Configure the same value on every gateway sharing the filesystem.

Checksum-derived ETags

--data-integrity-etag
VGW_DATA_INTEGRITY_ETAG=true

By default, regular uploads use MD5 ETags and completed multipart uploads use the conventional MD5-of-part-ETags form. This option instead gives new writes quoted, algorithm-qualified checksum ETags:

Write ETag form
Single PUT "ALGORITHM-<base64-checksum>"
Multipart part "CRC64NVME-<base64-part-checksum>"
Completed multipart, full-object checksum "ALGORITHM-<base64-whole-object-checksum>"
Completed multipart, composite checksum "ALGORITHM-<base64-composite-checksum>-<part-count>"

The client-selected checksum algorithm is used when available; otherwise the gateway computes CRC64NVME. Existing objects retain their stored ETags.

Avoiding the separate MD5 calculation may improve write performance when clients do not require MD5-compatible ETags. Benchmark representative transfers. Clients that parse ETags as hexadecimal MD5 digests or construct multipart ETags themselves may be incompatible. Configure the option consistently across gateway instances.

POSIX options not available on ScoutFS

The ScoutFS subcommand intentionally has a smaller option set:

POSIX option ScoutFS behavior
--sidecar, --nometa Unavailable; ScoutFS always stores metadata in xattrs
--io-buffer-size Unavailable; buffered fallback uses the POSIX backend default
--enable-odirect Unavailable; object reads and writes use buffered I/O
--disableotmp Unavailable; temporary-file selection is automatic
--disable-copy-file-range Unavailable; fallback copy selection is automatic
--disable-object-lock-file Unavailable; select --object-lock-mode=local explicitly

Related global options

Global option ScoutFS effect
--copy-object-threshold Rejects CopyObject and UploadPartCopy when the source exceeds the configured bytes; defaults to 5 GiB
--mp-max-parts Limits the number of parts accepted by multipart completion; defaults to 10,000
--disable-strict-bucket-names Relaxes frontend checks while the backend still rejects path-unsafe names; nonstandard names can remain absent from listings
--readonly Rejects all mutating S3 operations
--disable-acl Disables gateway ACL enforcement and ignores ACL request headers

Global options must precede scoutfs; backend options must follow it. See Global Options.

Object name mapping

Buckets map to top-level directories below the gateway root. Existing directories that pass bucket-name validation appear as buckets; top-level regular files are ignored.

Object keys map to directory components and files by splitting on /:

gateway root: /mnt/scoutfs/gateway
bucket:       mybucket
object:       2026/September/myobject

filesystem:   /mnt/scoutfs/gateway/mybucket/2026/September/myobject

A key ending in / represents a directory object and must have zero content length.

Glacier mode

ScoutFS can keep file metadata online while its data extents are offline in an archive. With --glacier, VersityGW maps that state to:

S3 operation ScoutFS behavior
GetObject Returns InvalidObjectState while any data blocks are offline
HeadObject Reports storage class GLACIER for offline data
HeadObject during staging Returns x-amz-restore: ongoing-request="true"
HeadObject offline and not staging Reports no active restore value
HeadObject online Returns ongoing-request="false" with a compatibility expiry date
ListObjects, ListObjectsV2 Report offline files with storage class GLACIER
RestoreObject Sets the ScoutFS batch stage-request flag

The reported expiry date does not control archive expiration; it is supplied for S3 client compatibility. VersityGW does not move bytes from archive storage itself. ScoutAM must observe the stage request and restore the extents.

Multipart extent-move optimization

UploadPart stores each part below:

<bucket>/.sgwtmp/multipart/<object-sha256>/<upload-id>/<part-number>

CompleteMultipartUpload first attempts to move each part's ScoutFS extents into the final object. Non-final parts and the current destination offset must be 4096-byte aligned; the last part may have any size. When alignment or an extent move prevents the optimization, the backend falls back to a normal file copy.

A successful extent move consumes the source part's extents and avoids a second read/write cycle. If a later part fails after a successful move, the upload can no longer be retried safely and is aborted during cleanup.

Limitations

  • All filesystem operations run as the gateway process identity. The gateway does not switch effective UID/GID per IAM request; ownership options affect new paths but not the process's access checks.
  • Out-of-band data changes do not update xattr metadata such as ETags, checksums, ACLs, policies, tags, retention, or legal holds.
  • A PUT cannot replace a directory at the same key and returns ExistingObjectIsDirectory. A file where a parent directory is required returns ObjectParentIsFile.
  • A trailing-slash directory object must be empty. Multipart creation for that key, or a non-empty PUT, returns DirectoryObjectContainsData.
  • Filesystem ENOSPC is returned as HTTP 507 InsufficientStorage; quota exhaustion is returned as QuotaExceeded.
  • The backend is not available on non-Linux or non-AMD64 builds.

Temporary files and atomic publication

In-flight PUT and UploadPart requests use temporary files. The inherited POSIX implementation attempts O_TMPFILE and automatically falls back to unique named files below .sgwtmp when unnamed files are unavailable. Unnamed files are reclaimed automatically when their descriptors close; named files receive best-effort cleanup and may remain after a crash.

The completed temporary file is atomically published into the object namespace. Concurrent unconditional writes therefore leave one complete uploaded value rather than a mixture of bytes; the last successful publication wins. Conditional writes additionally require --object-lock-mode=local and are atomic only within that gateway process.

Clone this wiki locally