The exit code of the test was a random number between 0 and 4,
which gave no guarantee that the right code from exit() will
actually reach the user.
Instead of testing a race between threads, make it possible to
call exit() in a chosen thread, and verify that it always works.
In addition, propagate the exit code from a forked process. That
use case was broken recently and has been fixed in the commit
titled 'Rework threads implementation'.
Most important differences from the old version:
- strip global process information from the thread struct into a
dedicated one,
- a parent is informed about the child death when the whole process
(the last thread) dies (not on each thread exit),
- all threads have the same parent (spawning thread is NOT the parent of
the spawned thread),
- a thread is able to wait on children created by another thread,
- a process is able to wait for exited children after execve,
- rewritten `waitid` implementation (no more gotos, supports __WCLONE
and friends flags),
- added option for syscall restarting, for now used only in `waitid`.
Additionally various bugfixes, cleanups and missing locks added.
Sometimes we need to temporarily stop IPC helper thread from receiving
more messages, e.g. when doing execve just before migrating exited (but
not yet waited for) children list.
* Make sure "stat.h" and "perm.h" are directly included where
necessary.
* Don't include "perm.h" inside "stat.h" but require it to be
included separately.
* Remove workarounds with __KERNEL__, __GLIBC__, defining pid_t
directly, and reversed include order (system headers before local
ones).
Instead of using S_I* flags, or hardcoded octal literals, use
helpers such as PERM_rwxrwxr_x. These are proposed in a Linux patch
by Ingo Molnar: https://lwn.net/Articles/696231/
Logging to file was broken, because the PAL file write operation
required the user to provide an absolute offset, and LibOS always
provided an offset of 0. This worked when logging to stdout, but
in case of a regular file, it kept overwriting the beginning of
file.
To fix that, we introduce a a special DkDebugLog call. This is a
better solution than tracking the file offset manually, because
the offset would need to be synchronized across different threads
and processes, and debug logs should be as simple as possible. At
the same time, we don't want PAL to provide a generic "append to
a file" mechanism, because it makes I/O less deterministic.
Pylint output was filtered so that many files with existing pylint
violations were allowed to stay broken.
I made sure all files pass pylint, but whitelisted some rules that
we commonly disable:
* missing docstrings: most of the code is tests/internal anyway
* invalid-name: too many violations, and we commonly use one- or
two-character names (like "a, b" or "t1, t2") which is
disallowed by this rule; we could tweak it and then fix
remaining violations such as camel-case or lowercase constants
* fixme: we leave TODOs as a matter of practice, same as in C
* high-level style rules like too-few-* and too-many-*,
no-self-use
Hopefully that will make using pylint less annoying, while also
catching serious issues (such as unused variables or imports).
Sometimes users may try to add `/dev` or root (`/`) mounts to the
manifest file, but these paths are already automatically mounted by
Graphene on startup. Previously, Graphene failed on assert in such
cases, now it prints a meaningful error message.
The manifest syntax stays exactly the same, including 0 and 1
integers to denote boolean values (this is done for ease of porting
and can be fixed in future commits). The only visible change is
surrounding strings in the manifest with quotes (requirement of
TOML). All manifests and Makefiles of our tests and example apps are
ported to the new TOML syntax. Documentation is updated.
The operation loops indefinitely on error. Instead, it should find
the first free FD, and then try allocating it.
In addition, the right error after exceeding the limit is EMFILE
(however, dup2() is still supposed to return EBADF if asking for
an out-of-range value, as checked by the dup201 LTP test).
They all lacked error checking and `wait_event` was completely broken:
it was reading from non-blocking pipe and treating EAGAIN as
successfully waited-for event.
Currently in Graphene, the type field in dirent struct is only updated
for parent directory and is set to DT_UNKNOWN type for its children.
But APIs like `sysconf(_SC_NPROCESSORS_CONF)` rely on type field to
identify the number of processors on the host by reading the number of
cpuX directories and ensuring their type is set to DT_DIR. This commit
addresses the issue by updating the dirent type appropriately even for
child directories or files.
Applications tend to use `/proc/cpuinfo` to get the `cpu cores`
and `physical id` for computing number of physical cores in a
socket. Currently `cpu cores` field is incorrectly implemented as
it is set to number of logical processors online and `physical id`
isn't implemented. This patch addresses both of these issues.
Previously, IPC_PORT_SERVER meant "listening port", and the actual
communication ports had several types. Only two of these types were
used for differentiation during IPC broadcast (direct-child and
direct-parent types). All other types denoted who is the remote party
this port connects to, but this info is superfluous. So this commit
replaces all these types with a generic IPC_PORT_CONNECTION, and
renames IPC_PORT_SERVER to a more familiar IPC_PORT_LISTENING.
As per linux man-pages, dirfd can be ignored if an absolute path is
provided when invoking *at() calls. But in Graphene many of the *at()
calls require a valid dirfd even when an absolute path is provided. If
not, the call fails. This commit fixes the issue by ignoring dirfd if
absolute filepath is provided.
Previously, process communication (channel between parent and newly
created child) was protected via TLS only during send/receive of the
checkpoint; after that the channel was downgraded from TLS to
plaintext. The reason for this downgrade is historical (IPC was
complicated, and we wanted to have at least some TLS at the time).
This commit fixes this issue: Graphene now always uses TLS on IPC.
Previously, Graphene assumed that if it was built with the DCAP
SGX driver or in-kernel SGX driver, then it should use DCAP/ECDSA
based attestation. In fact, the SGX driver has nothing to do with
the attestation scheme. This commit allows to use EPID based
attestation even when Graphene is built with the DCAP SGX driver.
- Remove old LTP bug workaround from Jenkins files
- Add pkg-config as an explicit dependency (it wasn't installed on
Ubuntu 16 and LTP won't build without it)
- Add short comments to tests that I've had time to investigate
- Collapse some test groups where we simply do not support the
feature being tested (e.g. a syscall) using a wildcard. Until we
implement it, we will not care about any new tests, and having
all the tests separately is just noise.
- Unskip some tests that seem to pass now
Keep only sections that differ from ltp.cfg, so that the file is
easier to maintain. Done automatically using
contrib/conf_subtract.py.
(As far as I can tell, there are no tests disabled in Linux and
enabled in Linux-SGX, only the other way around).