skyline

mirror of https://github.com/skyline-emu/skyline.git synced 2024-11-23 09:59:19 +01:00

Author	SHA1	Message	Date
Billy Laws	afef6c5123	Always populate all colour attachments This better follow the Vulkan spec, which doesn't mention anything about writes to OOB attachments, only those marked as unused.	2023-01-08 19:30:52 +00:00
Billy Laws	3571737392	Reset maxwell3d quick bind state before adding subpasses to executor If a submission happens during the call to addsubpass we could end up with invalid quick bind state, move this to to before to prevent that.	2023-01-08 19:30:52 +00:00
Billy Laws	3d31ade35f	Implement an alternative buffer path using direct memory importing By importing guest memory directly onto the host GPU we can avoid many of the complexities that occur with memory tracking as well as the heavy performance overhead in some situations. Since it's still desired to support the traditional buffer method, as it's faster in some cases and more widely supported, most of the exposed buffer methods have been split into two variants with just a small amount of shared code. While in most cases the code is simpler, one area with more complexity is handling CPU accesses that need to be sequenced, since we don't have any place we can easily apply writes to on the GPFIFO thread that wont also impact the buffer on the GPU, to solve this, when the GPU is actively using a buffer's contents, an interval list is used to keep track of any GPFIO-written regions on the CPU and any CPU reads to them will instead be directed to a shadow of the buffer with just those writes applied. Once the GPU has finished using buffer contents the shadow can then be removed as all writes will have been done by the GPU. The main caveat of this is that it requires tying host sync to guest sync, this can reduce performance in games which double buffer command buffers as it prevents us from fully saturating the CPU with the GPFIFO thread.	2023-01-08 19:30:52 +00:00
Billy Laws	b3f7e990cc	Allow for tying guest GPU sync operations to host GPU sync This is necessary for the upcoming direct buffer support, as in order to use guest buffers directly without trapping we need to recreate any guest GPU sync on the host GPU. This avoids the guest thinking work is done that isn't and overwriting in-use buffer contents.	2023-01-08 19:30:52 +00:00
Billy Laws	89c6fab1cb	Implement a way to check if the command record thread is idle Useful for debugging and testing	2023-01-08 19:30:52 +00:00
Billy Laws	c67f27e914	Add a setting to control the maximum number of accumulated GPU cmds This helps to keep the GPU fed when processing large command buffers which don't have any syncpoints to force a flush inbetween.	2023-01-08 19:30:52 +00:00
Billy Laws	77214a98dd	Add a setting to force maximum GPU clocks on KGSL devices	2023-01-08 19:30:52 +00:00
Billy Laws	83ecc33a77	Update adrenotools	2023-01-08 19:30:52 +00:00
Billy Laws	3ecaedd71e	Add adrenotools direct mapping support	2023-01-08 19:30:52 +00:00
Pablo	8846a85d3a	Stub some IPurchaseEventManager functions	2022-12-31 10:45:18 +00:00
PabloG02	80c0f8f04d	Implement full profile picture support Extends the profile picture stub into a full-fledged implementation with the ability for users to set their profile picture in settings while having the Skyline icon as the default profile picture.	2022-12-27 22:53:41 +05:30
PixelyIon	7a3d2e4a26	Start `KThread` TID from 1 rather than 0 HOS's TIDs are one-based rather than zero-based, certain titles such as Pokémon Arceus, Naruto Shippuden: Ultimate Ninja Storm 3, Splatoon 3, etc. use the TID being zero as a sentinel value but as we assigned this ID to our first thread prior it broke this logic which has now been fixed by this commit as it now matches HOS behavior.	2022-12-27 22:36:06 +05:30
Billy Laws	bab659587f	Use e1 sample count for blits	2022-12-22 18:05:45 +00:00
Billy Laws	516ece6b04	Calculate renderarea from attachment min size	2022-12-22 18:05:45 +00:00
Billy Laws	4a3cd69257	Populate graphics pipeline manager from cache at launch-time	2022-12-22 18:05:45 +00:00
Billy Laws	e9bcdd06eb	Introduce a pipeline cache manager for simple read/write cache accesses All writes are done async into a staging file, which is then merged into the main pipeline cache file at the time of the next launch. Upon encountering file corruption the cache can be trimmed up to the last-known-good entry to avoid any excessive loss of data from just one error.	2022-12-22 18:05:45 +00:00
Billy Laws	06bf1b38af	Introduce a pipeline state accessor that reads from a bundle	2022-12-22 18:05:45 +00:00
Billy Laws	7dd3a1db0f	Avoid InterconnectContext use in graphics PipelineManager We will soon move to a global pipeline manager instance, so it wont be possible to use InterconnectContext at pipeline-creation time anymore	2022-12-22 18:05:45 +00:00
Billy Laws	ffe7263848	Add quirk for 615 drivers with broken multithreaded compilation	2022-12-22 18:05:45 +00:00
Billy Laws	755f7c75af	Add pipeline (de)serialisation support to bundle See comments in code for details on the on-disk format.	2022-12-22 18:05:45 +00:00
Billy Laws	937eff392f	Switch execution-numbers to be globally unique tags This is required for making pipelines usable across channels without introducing caching bugs.	2022-12-22 18:05:45 +00:00
Billy Laws	072b8193a1	Implement thread pool based async pipeline compilation with futures By distributing the load of shader compiling onto multiple threads and then only waiting for completion until absolutely neccessary we can reduce compilation stutters significantly.	2022-12-22 18:05:45 +00:00
Billy Laws	186549748d	Implement HelperShader-local pipeline cache and use dynamic state Avoids the heavy overhead of the VK pipeline cache when we really only have a few bits of non-dynamic state	2022-12-22 18:05:45 +00:00
Billy Laws	9115b8cae8	Properly hash dynamic states in pipeline cache	2022-12-22 18:05:45 +00:00
Billy Laws	7c4b4765bf	Reduce thresholds for slot increase and buffer/texture fast readback	2022-12-22 18:05:45 +00:00
Billy Laws	f32ab1feff	Include BS thread pool library	2022-12-22 18:05:45 +00:00
Billy Laws	ce428af2e6	Use attachment formats rather than views in VK pipeline cache	2022-12-22 18:05:45 +00:00
Billy Laws	e849264028	Abstract out pipeline-compile-time GPU state accesses Introduces the base abstractions that will be used for pipeline caching, with a 'PipelineStateBundle' that can be (de)serialised to/from disk and an abstract accessor class to allow switching between creating disk-cached pipelines and fresh ones.	2022-12-22 18:05:45 +00:00
Billy Laws	2e96248fb6	Track RT format info in PackedPipelineState and move VK conv code there When caching pipelines we can't cache whole images, only their formats so refactor PackedPipelineState so that it can be used for pipeline creation, as opposed to passing in a list of attachments.	2022-12-22 18:05:45 +00:00
Billy Laws	bc7e1eb380	Split-out hash from ShaderBinary struct This isn't necessary for pipeline creation and creates some difficulty with pipeline caching.	2022-12-22 18:05:45 +00:00
Dima	de10ab1219	Stub SetConnectionConfirmationOption	2022-12-18 20:34:55 +00:00
Dima	f3b2b4317e	Stub some IPrepoService calls	2022-12-18 20:34:55 +00:00
Dima	efef67b92b	Stub some IAudioDevice calls	2022-12-18 20:34:55 +00:00
Dima	3a94bcf692	Fix ListOpenContextStoredUsers stub The problem is in StoreOpenContext wasn't storing any user, but ListOpenContextStoredUsers was writing default user (when it's not stored by StoreOpenContext)	2022-12-18 20:34:55 +00:00
TheASVigilante	3c5f8dd876	Fix small typo	2022-12-18 14:49:54 +00:00
lynxnb	6599c1dccf	Stub `GyroscopeZeroDriftMode` Related service calls are called in a loop by SM3DW. A variable tracking zero drift mode has been added to `npad_device`, but it's unused at the moment.	2022-12-10 14:59:44 +00:00
Dima	dcc3047ba8	Stub ErrorCommonArg	2022-12-10 14:58:20 +00:00
Dima	68253fe995	Stub mii:e/mii:u Needed for SSBU	2022-12-10 14:58:20 +00:00
Dima	69ee3cfc66	Stub DeleteDirectory Should allow deleting/rewriting saves in some games	2022-12-10 14:58:20 +00:00
Dima	bbd34ae7e7	Validate if entries are not empty before using Should fix saving problem in Baldur's Gate: Dark Alliance II at least	2022-12-10 14:58:20 +00:00
Dima	5f510d84d7	Stub IsVibrationPermitted	2022-12-10 14:58:20 +00:00
Dima	51d1f519af	Stub ListDisplays	2022-12-10 14:58:20 +00:00
Dima	a3866a3129	Stub LibraryAppletShop	2022-12-10 14:58:20 +00:00
Dima	1ebec7db82	Stub GetImageSize and LoadImage	2022-12-10 14:58:20 +00:00
Dima	52c4228ecf	Stub some friends service calls Needed for Diablo 3	2022-12-10 14:58:20 +00:00
Dima	ebcbc5b05b	Validate NpadId for ActivateVibrationDevice	2022-12-10 14:58:20 +00:00
Dima	4bdd033354	Stub SetRecordVolumeMuted	2022-12-10 14:58:20 +00:00
Dima	f6d95aae01	Stub GetCacheStorageSize	2022-12-10 14:58:20 +00:00
Dima	4ab8699cd4	Stub ImportServerPki	2022-12-10 14:58:20 +00:00
Dima	41cf4bb12d	Stub GetLanguageCode	2022-12-10 14:58:20 +00:00
Dima	3e078d54b6	Stub GetIdleTimeDetectionExtension	2022-12-10 14:58:20 +00:00
Dima	2311f777fc	Stub IsCpuOverclockEnabled	2022-12-10 14:58:20 +00:00
Dima	4601c28c28	Stub GetCurrentIpAddress	2022-12-10 14:58:20 +00:00
Dima	18e6a6c53c	Stub DeclareOpenOnlinePlaySession and DeclareCloseOnlinePlaySession	2022-12-10 14:58:20 +00:00
Dima	150c1370c2	Stub some IApplicationFunctions funcs	2022-12-10 14:58:20 +00:00
Dima	a6f3aa3062	Stub TrySelectUserWithoutInteraction and ListQualifiedUsers	2022-12-10 14:58:20 +00:00
Dima	5a9a2861df	Add TitleId TextView in App Dialog	2022-12-10 14:57:46 +00:00
Abandoned Cart	b08fcd7027	Favor a predefined "click" over system vibration	2022-12-10 14:57:33 +00:00
Abandoned Cart	cfd3bfecba	Add a rudimentary OSC button vibration setting	2022-12-10 14:57:33 +00:00
Billy Laws	7c802aea46	Mark vertex buffers as dirty on limit changes	2022-12-03 22:50:56 +00:00
Billy Laws	df19810c6c	Always set vertex stride for unbound buffers	2022-12-03 22:50:56 +00:00
Billy Laws	f4f658e3b7	Fix typo	2022-12-03 22:50:56 +00:00
Billy Laws	45b10ef776	Return whole mapping for shader code when end instrs aren't found	2022-12-03 22:50:56 +00:00
Billy Laws	d849875656	Only unlock GPU channel state on queue wait if it was previously locked	2022-12-03 22:50:56 +00:00
Billy Laws	a5e0a64adc	Switch patch error logs to debug	2022-12-03 22:50:56 +00:00
Billy Laws	af7c54297f	Cache staging buffer used for texture download	2022-12-03 22:50:56 +00:00
Billy Laws	8c5e6d2bb4	Update VKMA	2022-12-03 22:50:56 +00:00
Billy Laws	bba07fb101	Update for new hades	2022-12-03 22:50:56 +00:00
Billy Laws	a16383fd4b	Disable compute shaders on mali This will need to be debugged properly at some point but its fine for now.	2022-12-03 22:50:56 +00:00
Billy Laws	d69c6851f3	Update hades	2022-12-03 22:50:56 +00:00
Billy Laws	137d801843	Skip host1x HW emulation and effectively stub submission This was causing a bunch of logspam and isn't really needed as we will be using a HLE approach.	2022-12-03 22:50:56 +00:00
Billy Laws	579a2d9337	Add dynamic executor slot growth	2022-12-03 22:50:56 +00:00
Billy Laws	60169fce4c	Support 0-sized constant buffers	2022-12-03 22:50:56 +00:00
Billy Laws	b86dd99e1a	Align all SSBOs to 0x40 bytes Required by Adreno GPUs	2022-12-03 22:50:56 +00:00
Billy Laws	bfae292fb0	Make executor slot count setting exponential	2022-12-03 22:50:56 +00:00
Billy Laws	e0ae94be9d	Enable robustness1 Vulkan feature	2022-12-03 22:50:56 +00:00
Billy Laws	e8ef2d80af	CMake build file updates	2022-12-03 22:50:56 +00:00
Billy Laws	bf03f945ee	Implement the Kepler compute engine This can reuse a fair bit of the now-commonised Maxwell 3D code and mostly consists of compute-specific pipeline code which was deemed not suitable for being commonised (e.g. descriptor update code is somewhat duplicated). Of note is how compute lacks any active state at all de to its use of QMDs which bundle up all state into a single object in memory.	2022-12-03 22:50:56 +00:00
Billy Laws	4bc81f007f	Add some convinience helpers to compute engine regs	2022-12-03 22:50:56 +00:00
Billy Laws	4267a6af36	Add support for parsing and compiling compute shaders to the shader manager	2022-12-03 22:50:56 +00:00
Billy Laws	86dab65af4	Commonise maxwell3d state updater	2022-12-03 22:50:56 +00:00
Billy Laws	a0b81d54d6	Use pitch layout for linear RTs More likely to match in the texture cache when being sampled.	2022-12-03 22:50:56 +00:00
Billy Laws	ac85df7b7a	Start transition cache lookup with most recent one	2022-12-03 22:50:56 +00:00
Billy Laws	62c86b7690	Move maxwell3d to common constant buffer code	2022-12-03 22:50:56 +00:00
Billy Laws	8f0a6e78c5	Add Vulkan stride dynamic state and robustness support Fixes the waterfall in SMO by specifying vertex buffer bounds.	2022-12-03 22:50:56 +00:00
Billy Laws	23a7f70a8e	Commonise maxwell3d guest shader caching code	2022-12-03 22:50:56 +00:00
Billy Laws	6f6a312692	Commonise maxwell3d pipeline binding handling code A lot of pipeline code is difficult to commonise due to the inherent difference between compute and graphics pipelines, however the binding layout is shared so we can at least commonise that	2022-12-03 22:50:56 +00:00
Billy Laws	be8cbabd97	Commonise maxwell3d texture code This will be shared with the compute engine implementation.	2022-12-03 22:50:56 +00:00
Billy Laws	61e95c4b2c	Commonise maxwell3d sampler code This will be shared with the compute engine implementation, the only thing of note with this is that the binding register is now passed as a param since it is part of the compute QMD which can't be dirty tracked.	2022-12-03 22:50:56 +00:00
Billy Laws	7f93ec3df6	Commonise maxwell3d interconnect common code for use by other engines The compute engine will require most of this for basic functionality.	2022-12-03 22:50:56 +00:00
Billy Laws	281838fde1	Apply GPU readback hack to both buffers and textures And rename as appropriate.	2022-12-03 22:50:56 +00:00
Billy Laws	f358c4517e	Update edge credits	2022-12-03 22:50:56 +00:00
Billy Laws	eb00dc62f8	Implement support for 36 bit games by using split code/heap mappings Although rtld and IPC prevent TLS/IO and code from being above the 36-bit AS limit, nothing depends the heap being below it. We can take advantage of this by stealing as much AS as possible for code in the lower 36-bits.	2022-12-02 22:10:03 +00:00
Dima	e8e1b910c3	Add possibility to disable audio output	2022-12-02 00:33:28 +01:00
lynxnb	70109f8fbd	Work around invalid values in `CNTFRQ_EL0` register Exynos SoCs have a bug where the `CNTFRQ_EL0` register is either set to 0 or contain incoherent values. With this patch, the frequency value is loaded into a static variable and used instead of reading the register. The value will be initialised to the correct value for affected SoCs, while unaffected ones will use the value from the register.	2022-12-02 00:23:28 +01:00
lynxnb	54d0246ca6	Tweak `GpuDriverActivity` FAB padding	2022-11-28 00:06:07 +01:00
lynxnb	2e8d7b559c	Use the original view padding/margin when applying window insets Adding to the current view padding/margin values results in applying the insets over and over again as insets listeners can be called multiple times.	2022-11-28 00:04:39 +01:00
Billy Laws	b2384e83f5	Add prepo:a service	2022-11-25 16:26:00 +00:00
Billy Laws	736216a6f4	Stub OpenPatchDataStorageByCurrentProcess	2022-11-25 16:26:00 +00:00
Billy Laws	44033d7f8d	Adjust CalendarTime year to be relative to 0AD	2022-11-25 16:26:00 +00:00
Billy Laws	2ce2604421	Implement VFS file deletion	2022-11-25 16:26:00 +00:00
Billy Laws	6c968e0357	Fix GetEntryType IPC return type	2022-11-25 16:26:00 +00:00
lynxnb	ec220c8ea9	Use an extended FAB in `GpuDriverActivity`	2022-11-23 19:49:42 +05:30
lynxnb	163f4f2014	Fix window insets handling when in landscape mode To avoid code duplication, insets handling has been moved to a separate interface.	2022-11-23 19:49:42 +05:30
lynxnb	ab6c5f4c50	Improve robustness of `KeyReader.import` * Close the input and output file streams before moving the output file to the final destination * Clean up the destination path before moving the new file * Introduce a `ImportResult` return value to differentiate between the possible causes of import errors * Display more meaningful error messages in the UI	2022-11-23 19:49:42 +05:30
lynxnb	38129d9dc3	Mark some strings as non-translatable	2022-11-23 19:49:42 +05:30
lynxnb	ee8c055641	Make `GpuDriverInstallResult` PascalCase	2022-11-23 19:49:42 +05:30
Billy Laws	7f1667de82	Avoid using trapping for frequently trapped shaders Fall back to hashing for every shader access as that ends up being faster than applying traps for every execution.	2022-11-19 12:49:05 +00:00
Billy Laws	06095918a9	Introduce per-channel sequence number for invalidation tracking For cases like shaders, which may be uploaded through I2M (which no longer causes an execution) we need a way to cause an invalidation on all writes	2022-11-19 12:49:05 +00:00
Billy Laws	97e3f7fd34	Increase max swapchain image count	2022-11-19 12:49:05 +00:00
Billy Laws	c49119f5ef	Fixup depth bounds register arguments	2022-11-19 12:49:05 +00:00
Billy Laws	db3c5c33c4	Clamp depth bounds into 0-1 range	2022-11-19 12:49:05 +00:00
Billy Laws	e1bbd521d9	Fix potential circular queue submission race If a producer thread was waiting for the queue to have free space and the consumer thread hadn't yet acquired the production mutex a deadlock could occur	2022-11-19 12:49:05 +00:00
Billy Laws	13baf2312f	Add a workaround for sampling BGRA textures with a swizzle	2022-11-19 12:49:05 +00:00
Billy Laws	13a96c5aba	Implement a helper shader for partial clears These are not natively supported by Vulkan, so use a helper shader and colorWriteMask for the same behaviour.	2022-11-19 12:49:05 +00:00
Billy Laws	ac0e225114	Use vkCmdBlit for texture copies when formats dont match	2022-11-19 12:49:05 +00:00
Billy Laws	c8fc8f84ec	Fallback to RGBA888 for unsupported swapchain formats as opposed to swizzle	2022-11-19 12:49:05 +00:00
Billy Laws	e0bc0d3a97	Avoid megabuffering buffers larger than the chunk size	2022-11-19 12:49:05 +00:00
Billy Laws	b6f49884b3	Use lower_bound to speedup texture hostMapping lookup	2022-11-19 12:49:05 +00:00
Billy Laws	e7fda28ac6	Skip over textures in cache which have been replaced with a layer/mip match	2022-11-19 12:49:05 +00:00
Billy Laws	88cc696c7f	Only use 2D array depth targets when depth > 1	2022-11-19 12:49:05 +00:00
Billy Laws	7fed971b2d	Take firstIndex into account when calculating index (quad) buffer size Without this we would miss any elements beyond indexCount in the index buffer and they would be filled with random garbage causing vertex bombs	2022-11-19 12:49:05 +00:00
Billy Laws	1f9de17e98	Begin command buffers asynchronously in command executor vkBeginCommandBuffer can take quite some time on adreno, move it to the cycle waiter thread where it won't block GPFIFO.	2022-11-19 12:49:05 +00:00
Billy Laws	4b3e906c22	Update cached buffer execution number when megabuffering	2022-11-19 12:49:05 +00:00
Billy Laws	3ae1e78544	Match mip layers and array layers in texture manager	2022-11-19 12:49:05 +00:00
Billy Laws	d502adb309	Avoid WRW hazard in subpass deps	2022-11-19 12:49:05 +00:00
Billy Laws	e9313cc291	Use view layer count over texture for attachments	2022-11-19 12:49:05 +00:00
Billy Laws	e65ca52d91	Avoid potential buffer copy race	2022-11-19 12:49:05 +00:00
Dima	720cfaafb6	Stub caps:su	2022-11-18 15:35:03 +00:00
Dima	74afca4aab	Stub caps:u	2022-11-18 15:35:03 +00:00
Dima	27ff1ae19b	Stub caps:c	2022-11-18 15:35:03 +00:00
Dima	ffb0546609	Stub caps:a	2022-11-18 15:35:03 +00:00
Dima	1c8736cb56	Stub IsLargeResourceAvailable	2022-11-18 12:52:25 +00:00
Dima	dcd9e4ff61	Stub SetIdleTimeDetectionExtension, SetAlbumImageTakenNotificationEnabled	2022-11-18 12:52:25 +00:00
Dima	60843269de	Stub GetBlockedUserListIds and UpdateUserPresence	2022-11-18 12:52:25 +00:00
Dima	2cdfc7640c	Stub GetPreviousProgramIndex	2022-11-18 12:52:25 +00:00
Dima	360306eb61	Stub GetAddOnContentListChangedEventWithProcessId	2022-11-18 12:52:25 +00:00
Dima	3d475ca122	Stub GetAccountId	2022-11-18 12:52:25 +00:00
Dima	0b452fe36b	Stub GetFriendList	2022-11-18 12:52:25 +00:00
Dima	cc37d2231d	Stub CheckFreeCommunicationPermission and IsFreeCommunicationAvailable	2022-11-18 12:52:25 +00:00
Dima	ec81c97fa9	Stub TryPopFromFriendInvitationStorageChannel	2022-11-18 12:52:25 +00:00
Dima	413f162cf2	Stub some account functions	2022-11-18 12:52:25 +00:00
lynxnb	b209ae8e90	Account for stick flat area when retrieving axes value from a `MotionEvent`	2022-11-17 21:54:15 +01:00
lynxnb	c966220bab	Zero-initialize axes history instead of using null values Use zero initialization for axes history instead of using null values. Fixes the first axis event after launching a game being completely ignored.	2022-11-17 21:54:15 +01:00
lynxnb	3a657c44cc	Don't ignore HAT axes input events Capture HAT axes events ourselves instead of relying on the android framework to turn them into KeyCodes. Fixes handling of DPAD button presses on most controllers.	2022-11-17 21:54:15 +01:00
lynxnb	f1ec771944	Fix inverted axis polarity In the case of axis value being zero, polarity would favor one side of the stick resulting in invalid values. Fix that by taking into account axis history when calculating polarity.	2022-11-17 21:54:15 +01:00
lynxnb	675e8dbb2e	Move input handling code to a dedicated class	2022-11-17 21:54:15 +01:00
lybxlpsv	18861d73a3	Set systemUiVisibility during onResume for A11	2022-11-17 14:04:57 +01:00
Dima	262ee28611	Stub some bsd functions Co-authored-by: Lunar-Pixel <83507264+Lunar-Pixel@users.noreply.github.com>	2022-11-15 16:24:33 +00:00
Dima	9afa8b881e	Stub nsd:u/nsd:a and sfdnsres services	2022-11-15 16:24:33 +00:00
Billy Laws	01e27bd2dd	Implement ldr:ro LoadModule	2022-11-15 16:23:40 +00:00
Billy Laws	e571066409	Stub ldr:ro IRoInterface Some games initialise this service on startup however don't actually use it. Add a simple stub to allow such games to boot.	2022-11-15 16:23:40 +00:00
Billy Laws	1fc2641746	Stub the web applet	2022-11-13 11:37:18 +00:00
Billy Laws	021f82ef08	Stub ListOpenContextStoredUsers	2022-11-13 11:35:40 +00:00
Billy Laws	e7bab27d85	Fixup nvdrv channel private memory allocation This was incorrectly allocated in words, rather than bytes, meaning that guest allocations could overwrite the private memory and break inline syncpt operations	2022-11-13 11:35:40 +00:00
Billy Laws	8b523fa1f0	Avoid inline syncpt increments sending OOB GpEntries In cases where no wfi is required, the space where the WFI commands would go needs to be zeroed out to avoid the GPU reading uninitialised memory.	2022-11-13 11:35:40 +00:00
Billy Laws	cd0b2636e5	Prevent truncation of big page start in GetVaRegions	2022-11-13 11:35:40 +00:00
Billy Laws	f650f32bf0	Avoid duplicating NvDrv buffer unmap code	2022-11-13 11:35:40 +00:00
Billy Laws	001064b7bf	Fix GraphicsBufferProducer recreation We need to use a shared_ptr to ensure that the present callback doesn't do any UAFs, also unlocks the GBP during presentation as if the queue is full a deadlock could a rise where the present callback wouldn't be able to run due to the (waiting) DequeueBuffer thread holding the lock.	2022-11-13 11:35:40 +00:00
Billy Laws	29e89a3950	Fix crashes when opening non-existent directories	2022-11-13 11:35:40 +00:00
Billy Laws	ec139b3027	Fixup CancelBuffer fence handling	2022-11-13 11:35:40 +00:00
Billy Laws	7f24c7b857	Store KMemory object ptrs in memory class to avoid linear-time unmap This is quite a horrible solution but fixing it properly would require a whole rewrite of how we handle memory.	2022-11-13 11:35:15 +00:00
lynxnb	cc71e7b56c	Run emulation in a separate process Exiting from emulation has always been a big issue for Skyline, with guest and host threads that would keep running in the background unless the app was manually killed. Running emulation in a separate process allows us to kill it when we are done, avoiding the need for complex exiting management code.	2022-11-11 11:49:33 +00:00
lynxnb	281562ccdb	Fix FABs ripple effect in `OnScreenEditActivity`	2022-11-09 23:07:23 +05:30
lynxnb	56f6f8a362	Reword the unsupported gpu drivers message The old message was being misinterpreted as if the device's gpu was not supported by the emulator. Reword that message to explicitly mention custom drivers.	2022-11-09 23:07:23 +05:30
lynxnb	e2a5da1d67	Fix `AppDialog` layout * Add a drag indicator at the top * Fix flex layout wrapping when buttons didn't fit on a single line * Fix BottomSheetDialog peek height too small on landscape orientation * General cleanup of the layout	2022-11-09 23:07:23 +05:30
lynxnb	4146261069	Create a unified style for section titles	2022-11-09 23:07:23 +05:30
lynxnb	86364a84a2	Allow content to be drawn behind the navigation bar	2022-11-09 23:07:23 +05:30
lynxnb	6a6e89f070	Make `BottomSheetDialog` go fullscreen when fully expanded	2022-11-09 23:07:23 +05:30
lynxnb	f93d3b78d3	Add a drag indicator element at the top of LicenceDialog A new `DragIndicatorView` had been introduced, which draws a small drag handle element. When used inside a `BottomSheetDialog`, this view will add a callback for hiding the indicator when the dialog is fully expanded.	2022-11-09 23:07:23 +05:30
lynxnb	6848e69638	Improve design consistency across the app Game images, buttons and dialogs now have a consistent corner radius, across all game list layouts.	2022-11-09 23:07:23 +05:30
lynxnb	f63fdf26c9	Partially revert 'Make all `Dialog`s use `@color/backgroundColor` as the background color' This partially reverts commit `36a1f2a2ec`.	2022-11-09 23:07:23 +05:30
lynxnb	dec04db647	`AppListItem` misc tweaks * Restore text marquee on all layouts * Text size and color tweaks * List layout image has round corners * Clean up unneeded attributes	2022-11-09 23:07:23 +05:30
lynxnb	5c76a57e6e	Use `NestedScrollView` for licence dialogs and minor layout tweaks	2022-11-09 23:07:23 +05:30
Billy Laws	388245789f	Restructure ConditionalVariableSignal to avoid potential deadlock Since InsertThread can block for paused threads, we need to ensure we unlock syncWaiterMutex when calling it.	2022-11-09 23:02:26 +05:30
PixelyIon	f4a8328cef	Implement Symbol Hooking Symbol hooking is required for HLE implementations of certain features in the future such as `nvdec` and for more in-depth debugging of games as we can inspect them on a SDK function level which allows us to debug issues far more easily.	2022-11-07 23:56:22 +05:30
PixelyIon	8892eb08e6	Fix `MoveRegister` to clear when value is 0 The register wouldn't be cleared with a `MOVZ` when a value was zero due to the condition for writing an instruction requiring the `offsetValue` to be non-zero.	2022-11-07 23:56:22 +05:30
Billy Laws	f7ab3abb86	Allow load balancing when waiting on condvars	2022-11-06 20:47:26 +00:00
german77	b6e2fb894c	service: bcat: Stub CreateDeliveryCacheStorageService	2022-11-06 20:39:41 +00:00
Billy Laws	80b65d5094	Update XDR names	2022-11-03 22:53:01 +00:00
Billy Laws	a940d6fd34	Update submodules	2022-11-02 17:46:07 +00:00
Billy Laws	026bb04386	Impl some more texture formats	2022-11-02 17:46:07 +00:00
Billy Laws	133f08ed14	Stash new register value before executing deferred draws/updates Since the register writes technically happen after the draw, issues can occur if they happen before: e.g. skyrim updates ctSelect and disables all RTs after a draw, but this would happen before it previously and crash the driver.	2022-11-02 17:46:07 +00:00
Billy Laws	c50852e546	Implement the draw(...)BeginEnd Maxwell3D draw registers Used by guest Vulkan games and nouveau.	2022-11-02 17:46:07 +00:00
Billy Laws	270ef3e0d2	Implement GPFIFO semaphore acquire operations	2022-11-02 17:46:07 +00:00
Billy Laws	2ce146e28f	Don't crash on the Grp0SetSubDevMask TertOp Used by Vulkan games to set the SLI mask, not applicable to the switch.	2022-11-02 17:46:07 +00:00
Billy Laws	1d83dadefb	Drop size restruction bypass for frequently synced buffers In cases where large buffers are updated every draw this could seriously increase memory usage beyond 3GB in the megabuffer.	2022-11-02 17:46:07 +00:00
Billy Laws	1088ed514c	Introduce texture usage system to ensure RPs are split when necessary Vulkan doesn't allow sampling a texture and using it as an RT in the same RP, by tracking the texture usage status and splitting RPs when this occurs we can avoid such potential sync errors.	2022-11-02 17:46:07 +00:00
Billy Laws	2dd4698441	Adjust texture matching hacks	2022-11-02 17:46:07 +00:00
Billy Laws	4f5c9047ef	Add some additional texture formats used by Vulkan games	2022-11-02 17:46:07 +00:00
Billy Laws	6a830dfac5	Use shader-compiler side {S,U}Scaled format emulation	2022-11-02 17:46:07 +00:00
Billy Laws	579fd04117	Fixup ReadTextureType shader compiler callback	2022-11-02 17:46:07 +00:00
Billy Laws	b04d18eba5	Add support for split mappings to I2M uploads Used by Super Mario Sunshine and other Vulkan games.	2022-11-02 17:46:07 +00:00
Billy Laws	db5e208379	Clear images even when aspects mismatch	2022-11-02 17:46:07 +00:00
Billy Laws	3c8df327f1	Fixup subpass barriers and flags	2022-11-02 17:46:07 +00:00
Billy Laws	5ab80901c6	Drop some debug code	2022-11-02 17:46:07 +00:00
Billy Laws	4de89c8839	GPU NEW MARGEBAC	2022-11-02 17:46:07 +00:00
Billy Laws	7670c83405	Ensure textures are clean before paging them out	2022-11-02 17:46:07 +00:00
Billy Laws	1a2351386d	Add u64 iova ctor	2022-11-02 17:46:07 +00:00
Billy Laws	93d43e0115	Fully fill in swizzle component mappings Avoids the rest being default initialised to identity, which would break the intended effect of them.	2022-11-02 17:46:07 +00:00
Billy Laws	37ff0ab814	Add buffer manager support for accelerated copies These will be sequenced on the GPU/CPU depending on what's optimal and avoid any serialisation	2022-11-02 17:46:07 +00:00
Billy Laws	cac287d9fd	Implement accelerated uploads/copies through buffer manager Previously, both I2M uploads and DMA copies would force GPU serialisation if they happened to hit a trap or were used to copy GPU dirty buffers. By using the buffer manager to implement them on the host GPU we can avoid such slowdowns entiely.	2022-11-02 17:46:07 +00:00
Billy Laws	c5ec484d9a	Avoid redundantly passing executor in ctors when it's already in ChannelCtx	2022-11-02 17:46:07 +00:00
Billy Laws	463394ba72	Pass correct size for XFB buffers	2022-11-02 17:46:07 +00:00
Billy Laws	bd976676f4	Fix SNorm vertex formats	2022-11-02 17:46:07 +00:00
Billy Laws	b74098570f	Zero-out unused XFB varyings before passing to hades	2022-11-02 17:46:07 +00:00
Billy Laws	22f3ba6b93	Mark XFB buffers as GPU dirty	2022-11-02 17:46:07 +00:00
Billy Laws	26aeeaecf5	Add constant buffer GPU write pipeline barrier	2022-11-02 17:46:07 +00:00
Billy Laws	0b5d9308c4	Be more careful about potentially-unneeded GPU->CPU syncs These can be especially expensive so should be avoided as much as possible.	2022-11-02 17:46:07 +00:00
Billy Laws	e6530e2386	Delete graphics_context F	2022-11-02 17:46:07 +00:00
Billy Laws	ac2e6c125b	Switch to Roboto for Korean font	2022-11-02 17:46:07 +00:00
Billy Laws	b24a8465da	Don't require depthClamp	2022-11-02 17:46:07 +00:00
Billy Laws	9055c98e09	Only enable debug/verbose logs in (rel)debug builds	2022-11-02 17:46:07 +00:00
Billy Laws	0ebdbcf0ff	Don't lock stateMutex when updating buffer cycle	2022-11-02 17:46:07 +00:00
Billy Laws	dd360b8f75	Pass correct wait semaphore array size to queue submit	2022-11-02 17:46:07 +00:00
Billy Laws	c78a4b9699	Fixup buffer recreation to avoid deadlock when waiting on srcs	2022-11-02 17:46:07 +00:00
Billy Laws	d236bfe454	Enable depthClamp VK device feature	2022-11-02 17:46:07 +00:00
Billy Laws	95d849e1f6	Check FenceCycle signalled flag immediately before waiting The lock release within the wait for submission means that another thread could end up signalling the cycle and then the VK wait still happen after when the lock has been reacquired.	2022-11-02 17:46:07 +00:00
Billy Laws	1a23b929a7	Avoid chaining cycles in buffer recreation This had a chance of creating circular chains which obviously caused issues, just do a wait instead for now.	2022-11-02 17:46:07 +00:00
Billy Laws	a15db9cb06	Update hades submodule	2022-11-02 17:46:07 +00:00
Billy Laws	cfc55e60b0	Add robin map submodule	2022-11-02 17:46:07 +00:00
Billy Laws	6c0f084aae	Introduce hack to ignore frequently read-back textures Readback can be especially slow on mobile due to the varying load pattern it creates which often prevents the CPU/GPU from clocking up. Since some games perform texture readback but don't actually use it for anything significant implement a hack to skip it and significantly improve performance in such cases.	2022-11-02 17:46:07 +00:00
Billy Laws	e45e7546c8	Redesign buffer megabuffering Due to the frequency at which is is called megabuffering performance is critical to the performance of the entire emulator, especially in high-drawcall-count scenarios. After the view redesign, megabuffering on a per-view level was no longer possible nor desirable, and thus megabuffering was modified to just copy for every usage of a view. This worked great at the time since there were other bottlenecks, however gpu-new has since removed almost all of them and megabuffering is now a major sore point. Fix this by megabuffering small chunks and storing them in a page-table like structure within the buffer, these chunks can be referenced by multiple views and will be smartly invalidated whenever the sequence number or execution number changes to avoid any sequencing issues. In addition to this, to help the case where almost the whole buffer is read every single frame across a set of multiple views, an optimisation to skip the chunked tracking and use one large single megabuffer allocation and one single memcpy has been introduced. This reduces the overall amount of time spent in memcpy since large memcpys are quicker.	2022-11-02 17:46:07 +00:00
Billy Laws	7ea9aa52f5	Speed up reported guest GPU time Avoids triggering DRS in games in cases where it wouldn't actually benefit anything due to being CPU bottlenecked.	2022-11-02 17:46:07 +00:00
Billy Laws	31c2fb7d7a	Fixup IDirectory read	2022-11-02 17:46:07 +00:00
Billy Laws	7491178a9e	Pass base array layer to texture views	2022-11-02 17:46:07 +00:00
Billy Laws	ff57d2fbbf	Enforce stronger format and weaker dimension texture compat checks Rather than using just bpb for format compat, additionally check that the exact component bit layout matches since many games end up reusing RTs for unrelated textures. The texture size requirements have also been weaked to only check the resulting layer size as opposed to width/height - this is somewhat hacky but it gets around the problem of blocklinear alignment.	2022-11-02 17:46:07 +00:00
Billy Laws	14af383238	Only allow submitting `swapchainImageCount` images for host present at a time Prevents situations where nothing would otherwise be waiting on the GPU and since presentation no longer blocks too many images would be submitted for presentation.	2022-11-02 17:46:07 +00:00
Billy Laws	bcd96ac77d	Fixup A8R8G8B8 TIC format mapping 8-bit formats are inverted in TICs compared to Vulkan	2022-11-02 17:46:07 +00:00
Billy Laws	90466b8830	Implement depth clamp rasterisation state Used in SMO for shadows.	2022-11-02 17:46:07 +00:00
Billy Laws	1cfc4278f9	Disable preserve buffer/texture attachment opt for now Causes several issues and crashes in Pokemon without an obvious cause.	2022-11-02 17:46:07 +00:00
Billy Laws	e483cf9634	Use shader memory mirror when reading guest shaders Avoids triggering any traps that may be present on the region	2022-11-02 17:46:07 +00:00
Billy Laws	f6e4328b5a	Ensure blit src/dst textures are attached as execution cycle dependencies Since they're not in the TIC pool they would otherwise be freed	2022-11-02 17:46:07 +00:00
Billy Laws	77a131df60	Support using in-app renderdoc API to capture individual executions	2022-11-02 17:46:07 +00:00
Billy Laws	576bc6f37e	Add CommandExecutor slot count setting	2022-11-02 17:46:07 +00:00
Billy Laws	1a0819fb76	Use semaphores for presentation engine frame synchronisation Avoids waits on the CPU which can be costly and confuse the scheduler, also reduces latency significantly.	2022-11-02 17:46:07 +00:00
Billy Laws	0670e0e0dc	Support using Vulkan semaphores with fence cycles In some cases like presentation, it may be possible to avoid waiting on the CPU by using a semaphore to indicate GPU completion. Due to the binary nature of Vulkan semaphores this requires a fair bit of code as we need to ensure semaphores are always unsignalled before they are waited on and signalled again. This is achieved with a special kind of chained cycle that can be added even after guest GPFIFO processing for a given cycle, the main cycle's semaphore can be waited and then the cycle for the wait attached to the main cycle and it will be waited on before signalling.	2022-11-02 17:46:07 +00:00
Billy Laws	5b72be88c3	Stub ldn:u service	2022-11-02 17:46:07 +00:00
Billy Laws	77d76ed05a	Batch contiguous GMMU ranges into one	2022-11-02 17:46:07 +00:00
Billy Laws	e52dbf202f	Pass more Maxwell3D registers into interconnect	2022-11-02 17:46:07 +00:00
Billy Laws	83c7ed314e	Setup KThread pthread handle in StartThread Avoids a race with starting the thread and the handle not being set yet	2022-11-02 17:46:07 +00:00
Billy Laws	9784ae23e9	Skip checking affinity before taking load-balance WaitScheduler path The affinity mask may be set after the wait has began	2022-11-02 17:46:07 +00:00
Billy Laws	ad3195e06f	Split out guest texture layer size calcs into a seperate func	2022-11-02 17:46:07 +00:00
Billy Laws	8fa83fdf13	Fix deswizzling non-pow2 block size formats We need to use DivideCeil to avoid rounding off part of the texture. Fixes texture in Nier Automata: Game of the YoRHa edition.	2022-11-02 17:46:07 +00:00
Billy Laws	27de42f8df	Use surfaceClip as a hint for the underlying rendertarget size TIC sizes may not be aligned to block linear dimensions whereas RT sizes are and then limited by the surface clip. By using this to determine surface size we are more likely to get a match in texture manager for any future usages.	2022-11-02 17:46:07 +00:00
Billy Laws	297597f697	Fix texture manager depth compat comparison	2022-11-02 17:46:07 +00:00
Billy Laws	500f817a28	Synchronize all non-matching textures back to host before recreation	2022-11-02 17:46:07 +00:00
Billy Laws	05581f2230	Remove now redundant buffer/texture/megabuffer manager locks They have been superseeded by the global channel lock	2022-11-02 17:46:07 +00:00
Billy Laws	f5a141a621	Add dirty resource operator*	2022-11-02 17:46:07 +00:00
Billy Laws	b72720e8db	Finish off transform feedback implementation	2022-11-02 17:46:07 +00:00
Billy Laws	36fd885b49	Pack all draw state into a struct to avoid std::function allocations	2022-11-02 17:46:07 +00:00
Billy Laws	b5d0060c3f	Only use scissor for clear rect when enabled	2022-11-02 17:46:07 +00:00
Billy Laws	f93df35e6c	Only set line width when wideLines feature is supported	2022-11-02 17:46:07 +00:00
Billy Laws	4cebdfc8d3	Pass texture and cbuf state into pipeline manager for hades callbacks	2022-11-02 17:46:07 +00:00
Billy Laws	9ce848d4e0	Implement descriptor update batching and push descriptors Batching helps to avoid the need to attach so many objects to the fence cycle, which ends up taking a fair bit of time due to the allocation required.	2022-11-02 17:46:07 +00:00
Billy Laws	62a165b51e	Reformat maxwell3d interconnect codebase	2022-11-02 17:46:07 +00:00
Billy Laws	3766be59e7	Zero out vertex attribute state when disabled to avoid creating redundant pipelines	2022-11-02 17:46:07 +00:00
Billy Laws	751e3356e1	Keep shader trap lock held for the duration of an execution Avoids constant relocking on the GPFIFO thread (~0.5% of total time)	2022-11-02 17:46:07 +00:00
Billy Laws	314a9bccbc	Allow megabuffering readonly SSBOs	2022-11-02 17:46:07 +00:00
Billy Laws	4c2db0ba01	Implement ReadCbufValue and ReadTextureType hades callbacks Used for bindless and BRX instruction emulation.	2022-11-02 17:46:07 +00:00
Billy Laws	2163f8cde6	Implement alpha test pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	c86ad638c4	Keep track of transform feedback varyings pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	7ad2d94345	Zero out blend state when disabled to avoid creating redundant pipelines	2022-11-02 17:46:07 +00:00
Billy Laws	4052a93051	Force non-pushdescriptors for blit helper shader	2022-11-02 17:46:07 +00:00
Billy Laws	cb2a8c6d24	Enable wideLines Vulkan feature	2022-11-02 17:46:07 +00:00
Billy Laws	a3369637a9	Don't entirely wipe out per-index TIC cache efter each execution Keep a copy of the old TIC entry and view even after purge caches and use the execution number to check validity instead, if that doesn't match then just memcmp can be used as opposed to a full hash and map lookup.	2022-11-02 17:46:07 +00:00
Billy Laws	98c0cc3e7f	Impl preserve attached buffers/textures to avoid GPFIFO lock thrashing When profiling SMO, it became obvious that the constant locking of textures and buffers in SyncDescriptors took up a large amount of CPU time (3-5%), a precious resource in intensive areas like Metro. This commit implements somewhat of a workaround to avoid constant relocking, if a buffer is frequently attached on the GPU and almost never used on the CPU we can keep the lock held between executions. Of course it's not that simple though, if the guest tries to lock a texture for the first time which has already been locked as preserve on the GPFIFO we need to avoid a deadlock. This is acheived through a combination of two things: first we periodically clear the locked attachments every 2*SlotCount submissions, preventing a complete deadlock on the CPU (just a long wait instead) and meaning that the next time the resource is attached on the GPU it will not be marked for preservation due to having been locked on the guest before; second, we always need to unlock everything when the GPU thread runs out of work, as the perioding clearing will not execute in this case which would otherwise leave the textures locked on the GPFIFO thread forever (if guest was waiting on a lock to submit work). It should be noted that we don't clear preserve attached resources in the latter scenario, only unlock them and then relock when more work is available.	2022-11-02 17:46:07 +00:00
Billy Laws	0e8ccf1e99	Use memory_order_release for new descriptor set allocations	2022-11-02 17:46:07 +00:00
Billy Laws	34db5097da	Avoid using a shared_ptr reference to cycle for command buffer submission Somehow without this we can sometimes get crashes during appending to circular queue	2022-11-02 17:46:07 +00:00
Billy Laws	0428e8c7da	Support forcing regular descriptor sets in VK pipeline cache	2022-11-02 17:46:07 +00:00
Billy Laws	0f394d516b	Lock FenceCycle inbetween waiting on chained cycles and checking signalled Avoids one race where we would end up hogging all the locks of chained cycles and ourself when waiting for submission of previous cycles and prevent any forward progress due to another thread locking one of the chained cycles.	2022-11-02 17:46:07 +00:00
Billy Laws	a015fe753d	Only write npad controllerInfo entry on the HID thread if it is valid	2022-11-02 17:46:07 +00:00
Billy Laws	b5446846f7	Stub IsSixAxisSensorAtRest	2022-11-02 17:46:07 +00:00
Billy Laws	6719572b3b	Keep track of how often textures/buffers are locked on the CPU For the upcoming preserve attachment optimisation, which will keep buffers/textures locked on the GPU between executions, we don't want to preserve any which are frequently locked on the CPU as that would result in lots of needless waiting for a resource to be unlocked by the GPU when it occasionally frees all preserve attachments when it could have been done much sooner. By checking if a resource has ever been locked on the CPU and using that to choose whether we preserve it we can avoid such waiting.	2022-11-02 17:46:07 +00:00
Billy Laws	993ffb56f4	Avoid waiting on texture/buffer fence with trapMutex locked This could cause deadlocks under certain circumstances by preventing the GPFIFO from making any progress while also waiting on it at the same time.	2022-11-02 17:46:07 +00:00
Billy Laws	3e8bd26978	Add a global gm20b channel lock Allowing for parallel execution of channels never really benefitted many games and prevented optimisations such as keeping frequently used resources always locked to avoid the constant overhead of locking on the hot path.	2022-11-02 17:46:07 +00:00
Billy Laws	57a4699bd1	Add IOCTL trace events	2022-11-02 17:46:07 +00:00
Billy Laws	7861968c05	Fix memory::Buffer move constructor	2022-11-02 17:46:07 +00:00
Billy Laws	ef0ae30667	Implement Maxwell3D texture pool management and view creation Ontop of the TIC cache from previous code a simple index based lookup has been added which vastly speeds things up by avoding the need to hash the TIC structure every time.	2022-11-02 17:46:07 +00:00
Billy Laws	5542459c75	Use a SpinLock for guest shader code cache trap mutex	2022-11-02 17:46:07 +00:00
Billy Laws	3e12cde4d5	Make active Vulkan pipeline public	2022-11-02 17:46:07 +00:00
Billy Laws	2556966ec5	Don't attach textures to the active cycle in AttachTexture	2022-11-02 17:46:07 +00:00
Billy Laws	7dc3dde815	Introduce support for waiting for submission to FenceCycle Introducing async record resulted in breaking the assumption that any work submitted through command scheduler would be submitted in order with graphics submits. Since async record now unlocks the texture before it's submitted a seperate mechanism is needed to ensure ordering of submits. This is achieved by building support into fence cycle itself, with a conditional variable that is waited on for submission before any fence waits occur.	2022-11-02 17:46:07 +00:00
Billy Laws	54b85583ae	Fix layout transition in Texture::CopyFrom	2022-11-02 17:46:07 +00:00
Billy Laws	0f7c04ffb4	Use target format bpb when calculating linear mip level size	2022-11-02 17:46:07 +00:00
Billy Laws	849184452c	Add function to check if any guest texture mappings are unmapped	2022-11-02 17:46:07 +00:00
Billy Laws	98cb94ca6c	Bind an empty uniform buffer in place of unbound constant buffers	2022-11-02 17:46:07 +00:00
Billy Laws	55b85d0691	Implement combined image samplers and make descriptor code common between quick/normal updates	2022-11-02 17:46:07 +00:00
Billy Laws	ccf2d59351	Fixup input rate reading from packed pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	13970a5644	Refresh pipeline cached storage buffer bindings after each execution	2022-11-02 17:46:07 +00:00
Billy Laws	6bb2853ca0	Keep track of combined image samplers for quick bind	2022-11-02 17:46:07 +00:00
Billy Laws	040db37a28	Fix descriptor copies to be one per descriptor type	2022-11-02 17:46:07 +00:00
Billy Laws	d482e0ea98	Fix shader stage iteration to not miss the pixel stage	2022-11-02 17:46:07 +00:00
Billy Laws	0b808cc22b	Use Sint/Uint attribute type in place of Sscaled/Uscaled Scaled formats are not supported by any mobile GPUs.	2022-11-02 17:46:07 +00:00
Billy Laws	92ce220d3a	Ignore constant buffer selector size for updates Clustertruck performs a giant (0x3000 byte) update with a selector size of only 0x500, without this the update would only partially go through	2022-11-02 17:46:07 +00:00
Billy Laws	33f16ca26e	Handle unmapped blocks in CachedMappedBufferView	2022-11-02 17:46:07 +00:00
Billy Laws	5c0e4a839d	Fix SW BC2 decoding pitch	2022-11-02 17:46:07 +00:00
Billy Laws	60863fa162	Make viewports fallback to viewport 0 when their dimensions are invalid OpenGL games use a zero width/height for viewports 1-15 when they're unused	2022-11-02 17:46:07 +00:00
Billy Laws	586f872655	Update indexed quad conversion for new API	2022-11-02 17:46:07 +00:00
Billy Laws	498b4966d3	Avoid crashing on unmapped buffers Just log a warning instead	2022-11-02 17:46:07 +00:00
Billy Laws	9f2b20443b	Implement Maxwell3D draws	2022-11-02 17:46:07 +00:00
Billy Laws	5020478ace	Always rebind pipeline if it has changed from the previous draw	2022-11-02 17:46:07 +00:00
Billy Laws	af1b4ca4f8	Skip clears if attachments are invalid	2022-11-02 17:46:07 +00:00
Billy Laws	01a5e95ce1	Implement non-indexed quad conversion support in new GPU Uses memory::Buffer directly to allow for cleaner recreation and allow for removing host-only buffers from Buffer in the future.	2022-11-02 17:46:07 +00:00
Billy Laws	d42814bdc1	Add callback for when non-Maxwell3D engines alter pipeline state This is required so that Maxwell3D knows to restore all dynamic state before the next draw	2022-11-02 17:46:07 +00:00
Billy Laws	7133c5d6b3	Drop exclusiveSubpass in favour of only ending RPs when attachments change Significantly helps Mali performance by avoiding creating an excessive number of RPs.	2022-11-02 17:46:07 +00:00
Billy Laws	a3f38c0cf7	Add perfetto tracepoints to async record	2022-11-02 17:46:07 +00:00
Billy Laws	6dfef095e8	Up default record slot count to 6	2022-11-02 17:46:07 +00:00
Billy Laws	7d3a117a6f	Use a spinlock for descriptor set allocator	2022-11-02 17:46:07 +00:00
Billy Laws	78ddd03d1f	Mark newly allocated descriptor slots as active	2022-11-02 17:46:07 +00:00
Billy Laws	0867c593be	Support binding pipelines in state updater	2022-11-02 17:46:07 +00:00
Billy Laws	f42a0df72c	Use ctSelect register for colour targets and impl null color attachments Some games have invalid colour attachments sandwiched inbetween valid ones, these need to be skipped in Vulkan with UNUSED_ATTACHMENT	2022-11-02 17:46:07 +00:00
Billy Laws	38aad21d29	Share single flag variable for Maxwell3D batch draw/constant buffer update Slightly cheaper	2022-11-02 17:46:07 +00:00
Billy Laws	bd7eee8e2b	Optimise GPFIFO command processing GPFIFO code is very high throughput due to the sheer number of commands used for rendering. Adjust some types and switch to a if statement with hints to slightly increase processing speed.	2022-11-02 17:46:07 +00:00
Billy Laws	2cdf6c1fe6	Add branch hint attributes to a couple branches	2022-11-02 17:46:07 +00:00
Billy Laws	b310b99bdc	Handle unmapped ranges in TranslateRange	2022-11-02 17:46:07 +00:00
Billy Laws	acfa58ea8c	State uopdater MERGEBACK desC	2022-11-02 17:46:07 +00:00
Billy Laws	68b6b20f78	Bunch of cleanup/bugfixes for initial pipeline desc set impl Cleans up code slightly, fixes keeping track of per-stage descriptor counts and fixes storage buffer caching + more random stuff.	2022-11-02 17:46:07 +00:00
Billy Laws	eb4a9bab11	Mark freshly allocated descriptor slots as active	2022-11-02 17:46:07 +00:00
Billy Laws	f3184cdff1	Reorder active descriptor set slots to end of list Speeds up allocations by significantly reducing the number of needed atomic ops per alloc. Could be optimised further in the future if needed.	2022-11-02 17:46:07 +00:00
Billy Laws	128b68d8b2	Avoid resetting command buffers manually it's implicit Somewhat costly on adreno, and with the one time submit flag the reset is implicit for the next beginCommandBuffers call.	2022-11-02 17:46:07 +00:00
Billy Laws	e1717ed811	Implement Maxwell samplers	2022-11-02 17:46:07 +00:00
Billy Laws	f1600f5ad0	Support allocating into spans in the linear allocator	2022-11-02 17:46:07 +00:00
Billy Laws	04cea9239f	Implement descriptor set updating through StateUpdater Will automatically resolve buffer views at record-time into descriptor write/copy structures then apply the write and bind the set.	2022-11-02 17:46:07 +00:00
Billy Laws	d174ca950b	Revert "Reset executor command buffers asynchronously" This reverts commit fc7956df4ff56fdb2afc4b2bb0bbca82196179ca.	2022-11-02 17:46:07 +00:00
Billy Laws	2bbe975ea7	Reset executor command buffers asynchronously This took a little while to do on qcom drivers, moving it to the cycle waiter thread gives a tiny speedup.	2022-11-02 17:46:07 +00:00
Billy Laws	054d32567d	Allow mutation of input data by callback in CircularQueue::AppendTranform	2022-11-02 17:46:07 +00:00
Billy Laws	7c9212743c	Implement asynchronous command recording Recording of command nodes into Vulkan command buffers is very easily parallelisable as it can effectively be treated as part of the GPU execution, which is inherently async. By moving it to a seperate thread we can shave off about 20% of GPFIFO execution time. It should be noted that the command scheduler command buffer infra is no longer used, since we need to record texture updates on the GPFIFO thread (while another slot is being recorded on the record thread) and then use the same command buffer on the record thread later. This ends up requiring a pool per slot, which is reasonable considering we only have four slots by default.	2022-11-02 17:46:07 +00:00
Billy Laws	a197dd2b28	Allow for creating signalled fence cycles	2022-11-02 17:46:07 +00:00
Billy Laws	542651232b	Add a mutex to allow preventing buffer recreation	2022-11-02 17:46:07 +00:00
Billy Laws	379b4f163d	Implement popping from CircularQueue	2022-11-02 17:46:07 +00:00
Billy Laws	6d9dc9c6fb	Implement some more of Draw Now performs VK state updates on execution and takes advantage of quick descriptor binding.	2022-11-02 17:46:07 +00:00
Billy Laws	3b26f4f48a	Expose way to check inter-pipeline descriptor compatibility Allows users to skip/use quick descriptor sync even after switching pipelines.	2022-11-02 17:46:07 +00:00
Billy Laws	943a38e168	Implement StateUpdater for rapid recording of VK state updates Using command executor for each state individual update was found to be infeasible due to the shear number of state updates per draw and it relying on per-node heap allocations. Instead this commit takes advantage of each state update being used only once to implement a system of linearly-allocated state update commands that are linked together. After setting up all draw state with StateUpdateBuilder, the built StateUpdater can then be used in the execution phase to record all of the draw state into the command buffer with almost zero ovehead.	2022-11-02 17:46:07 +00:00
Billy Laws	7b4da52445	Add a fast binding sync path for when only one cbuf has changed SMO implements instanced draws by repeating the same draw just with a different constant buffer bound. Reduce the cost of this significantly by detecting such cases and instead of processing every descriptor, copy the previous descriptor set and update only the ones affected by the bound constant buffer. Credits to ripinperiperi for the initial idea and making me aware of how SMO does these draws	2022-11-02 17:46:07 +00:00
Billy Laws	89edd9b303	Reset megabuffer binding for disabled vertex buffers	2022-11-02 17:46:07 +00:00
Billy Laws	6a1615a104	Expose color and depth attachments to Draw	2022-11-02 17:46:07 +00:00
Billy Laws	aae957819e	Simplify BufferView locking by requiring buffer manager be locked Avoids the need to repeat the lookup after locking since recreations are impossible if buffer manager is locked.	2022-11-02 17:46:07 +00:00
Billy Laws	9449b52f36	Reduce minimum megabuffer alignment to 128 bytes	2022-11-02 17:46:07 +00:00
Billy Laws	b3cf9c40ba	Update megabuffer execution/sequence numbers after updating an allocation	2022-11-02 17:46:07 +00:00
Billy Laws	4b2b6fc6e9	Avoid calling SynchronizeGuest when attempting to megabuffer unless necessary	2022-11-02 17:46:07 +00:00
Billy Laws	e5919e84a1	Pipeline state if statment cleanups	2022-11-02 17:46:07 +00:00
Billy Laws	bf536aa168	Sync pipeline descriptors every draw	2022-11-02 17:46:07 +00:00
Billy Laws	9223d7f524	Fix descriptor initialisation order They need to be setup before the pipeline is created to avoid passing in garbage data.	2022-11-02 17:46:07 +00:00
Billy Laws	4652cc5a0a	Avoid parsing descriptors for disabled shader stages	2022-11-02 17:46:07 +00:00
Billy Laws	3456fb39fa	Fix pipeline to shader stage conversion when filling in shader infos The two vertex pipeline stages need to be both treated as a single stage, and all subsequent stages need to be offset by -1	2022-11-02 17:46:07 +00:00
Billy Laws	a9213debc7	Implement constant buffer reading	2022-11-02 17:46:07 +00:00
Billy Laws	afcfe8a7fa	Don't update scissor state >0 unless multiview is supported	2022-11-02 17:46:07 +00:00
Billy Laws	55d77b7eb0	Update user code for new megabuffering	2022-11-02 17:46:07 +00:00
Billy Laws	cc776ae395	Keep track of an 'execution number' in CommandExecutor Allows users to efficiently cache resources that are valid for only one execution without resorting to callbacks.	2022-11-02 17:46:07 +00:00
Billy Laws	99a34df4cc	Avoid trapping frequently synced buffers by using megabuffer copies When a buffer is trapped nearly every frame, the cost of trapping and synchronising its contents starts to quickly add up. By always using the megabuffer when this is the case, since megabuffer copies are done directly from the guest, we skip the need to synchronise/trap the backing.	2022-11-02 17:46:07 +00:00
Billy Laws	a24aec03a6	Rework per-view megabuffering to cache allocs in the buffer itself The original intention was to cache on the user side, but especially with shader constant buffers that's difficult and costly. Instead we can cache on the buffer side, with a page-table like structure to hold variable sized allocations indexed by the aligned view base address. This avoids most redundant copies from repeated use of the same buffer without updates inbetween.	2022-11-02 17:46:07 +00:00
Billy Laws	b810470601	Invalidate HLE macro state on macro updates	2022-11-02 17:46:07 +00:00
Billy Laws	2360ca24da	Implement constant buffer and storage buffer pipeline descriptor types	2022-11-02 17:46:07 +00:00
Billy Laws	25255b01c7	Keep track of more pipeline descriptor information This is needed in order to allow allocating per-draw descriptor arrays without recounting their lengths each time	2022-11-02 17:46:07 +00:00
Billy Laws	ad0275dbef	Expose active pipeline for access by Maxwell3D class	2022-11-02 17:46:07 +00:00
Billy Laws	6e22373b59	Add array support to AllocateUntracked	2022-11-02 17:46:07 +00:00
Billy Laws	388cff3353	Implement simple pipeline transition cache Avoids the need to hash PipelineState when we can guess the pipeline that will be used next. This could very easily be optimised in the future with generational, usage-based caching if necessary.	2022-11-02 17:46:07 +00:00
Billy Laws	302b2fcc3f	Force flush when dirty refresh returns true	2022-11-02 17:46:07 +00:00
Billy Laws	ec4ea5c5d7	Supply dispatcher manually for shader creation	2022-11-02 17:46:07 +00:00
Billy Laws	3404a3abdb	Implement macro HLE for instanced draw macros gm20b performs instanced draws by repeating draw methods for each instance, the code to detect this together with the cost of interpreting macros took up around 6% of GPFIFO time in Metro Kingdom. By detecting these specific macros and performing an instanced draw directly much of that cost can be avoided.	2022-11-02 17:46:07 +00:00
Billy Laws	cf0752f937	Use NCE memory tracking for guest shaders Prevents needing to hash them for every single pipeline state update, without this just hashing shaders takes up a significant amount of time.	2022-11-02 17:46:07 +00:00
Billy Laws	19a75c3f65	Bind all pipeline states to main pipeline dirty state	2022-11-02 17:46:07 +00:00
Billy Laws	a04d8fb5cf	Setup minimal viewport Vulkan pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	fe51db366b	Mark all dirty resources as dirty initially	2022-11-02 17:46:07 +00:00
Billy Laws	abfa5929f1	Treat vertex buffers with base addr 0 as disabled	2022-11-02 17:46:07 +00:00
Billy Laws	e71ca05f19	Avoid bitfields for signed enum types in PackedPipelineState	2022-11-02 17:46:07 +00:00
Billy Laws	2f2b615780	Add dynamic state support to VK graphics pipeline cache	2022-11-02 17:46:07 +00:00
Billy Laws	ad045058ee	Cleanup some redundant comments and includes in pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	9b05c9c0c3	Introduce a pipeline manager and partial pipeline object gpu-new will use a monolithic pipeline object for each pipeline to store state, keyed by the PackedPipelineState contents. This allows for a greater level of per-pipeline optimisations and a reduction in the overall number of lookups in a draw compared to the previous system.	2022-11-02 17:46:07 +00:00
Billy Laws	2c1b40c9a8	Add some missed pipeline state members used by Hades	2022-11-02 17:46:07 +00:00
Billy Laws	0c6fa22c6b	Introduce pipeline shader stage state Simple state that generates a hash that can be used in the packed state and spans over the binary for pipeline creation.	2022-11-02 17:46:07 +00:00
Billy Laws	a6bb716123	Move packed pipeline state to a seperate file	2022-11-02 17:46:07 +00:00
Billy Laws	4dcbf5c3a0	Drop the caching aspect of shader manager entirely Caching here was deemed unnecessary since it will be done implicitly by the pipeline cache and creates issues with the legacy attribute conversion pass. It now purely serves as a frontend for Hades.	2022-11-02 17:46:07 +00:00
Billy Laws	e77e4891dc	Tidy up Maxwell 3D regs a bit	2022-11-02 17:46:07 +00:00
Billy Laws	405d26fc22	Introduce Maxwell 3D shader state Simple state holder that hashes the stored shader and reads it into a buffer	2022-11-02 17:46:07 +00:00
Billy Laws	6865f0bdaf	Transition color blend state to packed pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	7049a521d2	Use Vulkan types directly in PackedPipelineState where possible	2022-11-02 17:46:07 +00:00
Billy Laws	effeb074b6	Move pipeline cache Key to cpp file and rename to PackedPipelineState The new name is more indicative of what the struct contains and how it will be used as more than just a key.	2022-11-02 17:46:07 +00:00
Billy Laws	e1512c91a0	Transition depth stencil state to pipeline cache key	2022-11-02 17:46:07 +00:00
Billy Laws	1f844e2c18	Transition rasterization state to pipeline cache key	2022-11-02 17:46:07 +00:00
Billy Laws	9a6efb091c	Transition tessellation state to pipeline cache key Also adds dirty tracking and removes it from direct state while we're at it. Since we no longer use Vulkan structs directly there's no benefit to it.	2022-11-02 17:46:07 +00:00
Billy Laws	ae5d419586	Transition input assembly state to pipeline cache key	2022-11-02 17:46:07 +00:00
Billy Laws	3f9161fb74	Transition vertex input state to pipeline cache key Also adds dirty tracking and removes it from direct state while we're at it. Since we no longer use Vulkan structs directly there's no benefit to it.	2022-11-02 17:46:07 +00:00
Billy Laws	ffe24aa075	Introduce a base pipeline cache key starting with RTs It was determined that a general purpose Vulkan pipeline cache isn't viable for the significant performance reqs of Draw(), by using a Maxwell 3D specific key we can shrink state significantly more than if we used Vulkan structs.	2022-11-02 17:46:07 +00:00
Billy Laws	1e8f7d7fcb	Implement dynamic pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	a94040ac7d	Add default cases to enum conversions where necessary	2022-11-02 17:46:07 +00:00
Billy Laws	6e55d4dcf4	Implement color blend pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	cb11662ea5	Implement depth stencil pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	2484a2d6b5	Add A1R5G5B5Unorm and S8Uint formats	2022-11-02 17:46:07 +00:00
Billy Laws	690e96bce0	Implement rasterization pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	90cd6adb91	Implement tessellation pipeline state	2022-11-02 17:46:07 +00:00
Billy Laws	3649d4c779	Adapt Maxwell 3D engine to new interconnect code Removes all usage of graphics_context.h from the codebase, exclusively using the new interconnect and its dirty tracking system. While porting the code a number of bugs were discovered such as not respecting the base instance or primitive type override, which have all been fixed. Currently only clears and constant buffer updates are implemented but due to the dirty state system allowing register handling on the interconnect end there shouldn't end up being many more changes.	2022-11-02 17:46:07 +00:00
Billy Laws	ef11900a39	Introduce reworked Maxwell 3D core interconnect This mainly distributes operations down to activeState and pipelineState, aside from clears which are implemented in-place. The exposed interface is much reduced as opposed to the previous GraphicsContext system due to the newly introduced dirty system, this should hopefully make the code more maintainable and keep actual rendering operations seperate from primitive restart state or whatever. Currently draws are unimplemented and the only full implemented things are clears and constant buffer operations.	2022-11-02 17:46:07 +00:00
Billy Laws	37b821a4dc	Introduce Maxwell 3D interconnect active state Active state encapsulates all state that isn't part of a pipeline and can be set dynamically with Vulkan calls. This includes both dynamic state like stencil faces, and command buffer state like vertex buffer bindings. Simililarly to the last commit, the main goal of this is to reduce the number of redundant work done per draw by employing dirty state as much as possible. Without using dirty state for this every active state operation would need to be performed every draw, which gets very expensive when things like buffer lookups end up being reqiored. Code has also been heavily cleaned up as is described in the previous commit.	2022-11-02 17:46:07 +00:00
Billy Laws	5fdda78073	Introduce Maxwell 3D interconnect pipeline state The main goal of this is to reduce the number of redundant lookups and work done per draw as much as possible, this is mainly achived through heavy used of dirty tracking though other optimisations like heavily using the linear allocator are also in play. In addition to the goal of performance, the code has been cleaned up and abstracted significantly from its state in graphics_context, hopefully making the GPU interconnect code much more maintainable in the future and reducing the boilerplace needed to add even simple functionality. This commit includes partial pipeline state, enough for implementing clears + a slight bit extra.	2022-11-02 17:46:07 +00:00
Billy Laws	21f5611231	Rewrite constant buffer interconnect code Adepted from the previous code to use dirty state tracking. The cache has also been removed since with the new buffer view and GMMU optimisations it actually ended up slowing lookups down, another result of the buffer view optimisations is that raw pointers are no longer used for buffer views since destruction is now much cheaper.	2022-11-02 17:46:07 +00:00
Billy Laws	d1e7bbc1d8	Introduce common code for Maxwell 3D interconnect rewrite This common code will be used across the entirety of the 3D rewrite, it also includes a stub for StateUpdateBuilder, which will be used by active state code to apply state updates.	2022-11-02 17:46:07 +00:00
Billy Laws	a6c49115f9	Rewrite all Maxwell 3D registers up to clears to match Nvidia docs All the names are directly translated from Nvidia docs, with minimal conversions to enums/structs when appropriate. Not all registers have been rewritten, only those that are needed to implement clears and dynamic state, the rest will be added as they are used in the GPU rework.	2022-11-02 17:46:07 +00:00
Billy Laws	d7eab40f1c	Introduce resource based dirty tracking infrastructure This will be heavily used by the upcoming GPU rework. It provides an intuitive way to track dirtiness based on using the underlying pointers of objects, as opposed to other methods which often need an enum entry per dirty state and don't support overlaps. Wrappers for dirty state objects are also provided to abstract as much of the dirty tracking as possible from user code. The pointer based mechanism also serves to avoid having to handle dirty bindings on the user side of the dirty resources, allowing them to bind things internally instead.	2022-11-02 17:46:07 +00:00
Billy Laws	8471ab754d	Introduce a spin lock for resources locked at a very high frequency Constant buffer updates result in a barrage of std::mutex calls that take a lot of time even under no contention (around 5%). Using a custom spinlock in cases like these allows inlining locking code reducing the cost of locks under no contention to almost 0.	2022-11-02 17:46:07 +00:00
Billy Laws	d810619203	Drop 3D engine method calling fast path in GPFIFO This ended up actually turning out to be a slow path when Maxwell 3D method handling code was inlined.	2022-11-02 17:46:07 +00:00
Billy Laws	ded02e3eac	Small engine.h fixups	2022-11-02 17:46:07 +00:00
Billy Laws	38ba963311	Drop usage of unique_ptr for Maxwell3D Since graphics context is being replaced and split into cpp files there will no longer be any circular includes that previously prevented this.	2022-11-02 17:46:07 +00:00
Billy Laws	90db743c56	Source AsGpu GMMU page sizes from GMMU class	2022-11-02 17:46:07 +00:00
Billy Laws	e72fe02c15	Add inline fast-path for `Buffer::FindOrCreate()` This can be inlined by the compiler much easier which helps perf a fair bit due to the number of times buffers are looked up, also avoids the need for small vector construction that was done in the previous fast-path.	2022-11-02 17:46:07 +00:00
Billy Laws	49478e178a	Avoid redundantly syncing buffers before every Write in an execution This isn't a guarantee provided by actual HW so we don't need to provide it either, the sync can be skipped once the buffer already been synced at least once within the execution.	2022-11-02 17:46:07 +00:00
Billy Laws	f7a726e452	Allow attempting to write to buffers without passing a GPU copy callback Constructing the GPU copy callback in `ConstantBuffers::Load()` ended up taking a fair amount of time despite it almost never being used in practice. By making it optional it can be skipped most of the time and only done when it's actually neccessary by calling `Write()` again if the initial call returned true.	2022-11-02 17:46:07 +00:00
Billy Laws	5dca5cc10e	Redesign buffer view infra to remarkably reduce creation overhead Buffer views creation was a significant pain point, requiring several layers of caching to reduce the number of creations that introduced a lot of complexity. By reworking delegates to be per-buffer rather than per-view and then linearly allocating delegates (without ever freeing) views can be reduced to just {delegatePtr, offset, size}, avoiding the need for any allocations or set operations in GetView. The one difficulty with this is the need to support buffer recreation, which is achived by allowing delegates to be chained - during recreation all source buffers have their delegates modified to point to the newly created buffer's delegate. Upon accessing a view with such a chained delegate the view will be modified to point directly to the end delegate with offset being updated accordingly, skipping the need to traverse the chain for future accesses.	2022-11-02 17:46:07 +00:00
Billy Laws	09f376e500	Add const accessors to OffsetMember	2022-11-02 17:46:07 +00:00
Billy Laws	64a9db2e82	Introduce `MergeInto` helper for simplified construction of arrays of structs In the upcoming GPU code each state member will hold a reference to its corresponding Maxwell 3D regs, this helper is needed to allow easy transformation from the the main 3D register struct into them. Example: ```c++ struct Regs { std::array<View, 10> viewRegs; u32 enable; } regs; struct ViewState { const View &view; const u32 &enable; size_t index; }; std::array<ViewState, 10> viewStates{MergeInto<ViewState, 10>(regs.viewRegs, regs.enable, IncrementingT{}) ```	2022-11-02 17:46:07 +00:00
Billy Laws	2c682f19a6	Add untracked linear allocator emplace/allocate functions Useful for cases where allocations are guaranteed to be unused by the time `Reset()` is called and calling `Free()` would be difficult or add extra performance cost due to how the allocation is used.	2022-11-02 17:46:07 +00:00
Billy Laws	6359852652	Introduce page size constants and replace all usages of PAGE_SIZE Avoids using macros and results in code which looks slightly cleaner.	2022-11-02 17:46:07 +00:00
Billy Laws	30ec844a1b	Use GPFIFO pushbuffer contents in-place if possible The memcpy within `Read()` was taking up a fair amount of time, avoid this by using the mapped range in-place when the mapping isn't split.	2022-11-02 17:46:07 +00:00
Billy Laws	be825b7aad	Utilise SegmentTable for rapid FlatMemoryManager lookups In some games performing the binary search in `TranslateRange()` ended up taking a fairly large (~8%) proportion of GPFIFO time. By using a segment table for O(1) lookups this is reduced to <2% for non-split mappings at the cost of slightly increased memory usage (2GiB in the absolute worse case but more like 50MiB in real world situations). In addition to adapting `TranslateRange()` to use the segment table, a new function `LookupBlock()` for cases where only a single mapping would ever be looked up so the small_vector handling and fallback paths can be skipped and the entire lookup be inlined.	2022-11-02 17:46:07 +00:00
Morph	4ea0b0e1e5	fssrv: IFileSystemProxy: Implement OpenReadOnlySaveDataFileSystem Forward this function to OpenSaveDataFileSystem for now. A proper implementation should wrap the underlying filesystem with nn::fs::ReadOnlyFileSystem.	2022-10-30 20:04:40 +00:00
Narr the Reg	25b9bb00fd	service: hid: Properly clear and set npad devices	2022-10-30 15:52:03 +00:00
german77	cf95cfb056	service: hid: Stub SetPalmaBoostMode	2022-10-30 15:52:03 +00:00
Narr the Reg	4da934579c	service: hid: Set the correct maxEntry value and signal on acquire event handle	2022-10-30 15:52:03 +00:00
Dima	baa6b5d5ea	check if NpadId is valid when update	2022-10-30 15:51:45 +00:00
Dima	a409f30e91	add GetAvailableLanguageCodeCount for both lists	2022-10-30 15:51:29 +00:00
Dima	51ce3f7c3c	Stub IClient::Poll	2022-10-29 19:52:50 +05:30
Billy Laws	1846f533bc	Add credits PreferenceCategory	2022-10-25 21:40:28 +01:00
Abandoned Cart	267d25b5a5	Prevent a false positive `SecurityException` for DocumentsProvider	2022-10-25 20:17:18 +05:30
lynxnb	1ae36cea24	Update `ngword2` archive with the correct content	2022-10-25 00:40:46 +02:00
Billy Laws	160c2f3457	Add Ko-Fi credits to settings	2022-10-23 21:14:39 +01:00
PixelyIon	128ea33073	Print NPDM + NACP metadata for title determination We determined that printing NPDM + NACP metadata is a significantly better way to determine what title is running rather than printing the filename.	2022-10-23 20:20:44 +05:30
PixelyIon	a0539a3edb	Trace Scheduler Preemption/Yield in Perfetto	2022-10-22 17:37:03 +05:30
PixelyIon	c874907eb5	Log and flush inside `KProcess::Kill` We want to know when the `KProcess` is being killed and flushing log during it is important since it can often result in hangs due to joining not working correctly.	2022-10-22 17:17:04 +05:30
PixelyIon	597a6ff31d	Wait on slot to be freed in `GraphicBufferProducer::DequeueBuffer` We currently don't wait on a slot to be freed if none are free, this worked prior to async presentation as GBP's slots wouldn't change their state until other commands were called but now slots can be held by the presentation engine. As a result, we now have to wait on the presentation engine to free up slots. This commit also fixes the behavior of the `async` flag in `DequeueBuffer` as it was treated as a non-blocking flag but isn't supposed to do anything on HOS.	2022-10-22 17:15:33 +05:30
lynxnb	cdb2b85d6c	Stub `ngword` and `ngword2` Co-Authored-By: Timotej Leginus <35149140+timleg002@users.noreply.github.com>	2022-10-18 20:54:57 +01:00
lynxnb	a7be4fd1e1	Stub `pl::RequestLoad` and `pl::GetSharedFontInOrderOfPriority` Co-Authored-By: Timotej Leginus <35149140+timleg002@users.noreply.github.com>	2022-10-18 20:54:57 +01:00
lynxnb	782f9e37ee	Add a system region setting Needed for games such as AC:NH. The `Auto` option automatically selects a region based on the currently selected system language. Co-Authored-By: Timotej Leginus <35149140+timleg002@users.noreply.github.com>	2022-10-18 20:54:57 +01:00
lynxnb	d4800d13b8	Stub `hid::ActivateMouse` and `hid::ActivateKeyboard` Co-Authored-By: Timotej Leginus <35149140+timleg002@users.noreply.github.com>	2022-10-18 20:54:57 +01:00
lynxnb	45830633eb	Stub `am::SetTerminateResult`	2022-10-18 20:54:57 +01:00
lynxnb	bc016aff47	Make the vulkan validation layer toggleable via setting As part of this commit, a new preference category for debug settings is being introduced. All future settings only relevant for debugging purposes will be put there. The category is hidden on release builds.	2022-10-18 19:47:23 +02:00
lynxnb	5cf14e45e1	Enable frame throttling when triple buffering is disabled	2022-10-18 19:27:14 +02:00
lynxnb	6b76c61cd1	Introduce a `releasedebug` build variant	2022-10-17 18:39:32 +02:00
lynxnb	b17364bb92	Introduce a `dev` app flavor for side-by-side installation	2022-10-01 13:01:46 +02:00
Billy Laws	8d1026d0cc	Reduce font shared memory size for compacted fonts	2022-09-22 21:51:13 +01:00
Billy Laws	25fb09800a	Compact fonts to only include necessary glyphs	2022-09-22 21:51:13 +01:00
Billy Laws	e31ed6a429	Increase font shared memory size for Noto fonts	2022-09-22 21:34:29 +01:00
Billy Laws	0d4893c448	Cleanup font magic generation code	2022-09-22 21:34:29 +01:00
Billy Laws	5c4cc3d51f	Fix font load order to match HOS Without this games would load an inappropriate font file for their target language.	2022-09-22 21:34:29 +01:00
Billy Laws	272bbf6cd2	Switch to Noto Sans fonts for shared fonts replacement Provides CJK characters, which the previous replacements lacked entirely.	2022-09-22 21:34:29 +01:00
lynxnb	54172322fe	Fix host synchronization for texture with a different guest format Host synchronization of a guest texture with a different guest format represents a valid use case where the host doesn't support the guest format and conversion to a host-compatible format must be performed. The issue is most evident on Mali GPUs, as they don't support BCn texture formats thus needing manual decoding before submission. It was disabled by mistake in a previous commit, this commit re-enables it.	2022-09-15 15:22:52 +02:00
TheASVigilante	a1ff4e1777	Implement OpenHardwareOpusDecoderEx and GetWorkBufferSizeEx Implements these 2 functions which were introduced in HOS 12.0.0. Fixes a crash in Xenoblade Chronicles 3.	2022-09-10 22:04:15 +01:00
lynxnb	34bd16426c	Fix quads index buffer conversion not accounting for first index Unindexed quad draws were broken when multiple draw calls were done on the same vertex buffer, with a non-zero `first` index. Indexed quad draws also suffered from the same issue, but was never encountered in games. This commit fixes both cases by accounting for the `first` drawn index when generating conversion index buffers.	2022-09-04 12:42:33 +02:00
Billy Laws	9af5df4bae	Increase IPC pointer buffer size Some services report a pointer buffer size > 0x1000, so up it as to not cause issues when they're implemented.	2022-09-02 23:14:05 +01:00
Billy Laws	51f4e7662e	Add support for the TIPC protocol introduced in HOS 12.0.0 TIPC is a much lighter layer ontop of the Horizon IPC system than CMIF and is used by SM in 12.0.0+. This implementation is slightly hacky since it doesn't really keep a seperation between the underlying kernel IPC stuff and userspace like CMIF/TIPC, this should be fixed eventually, probably together with an IPC dispatch rewrite to avoid the mess of frozen maps. Tested with Hentai Uni, which now crashes needing 'ldr:ro'.	2022-09-02 23:13:23 +01:00
Billy Laws	a40d7c78ad	Always recreate oboe stream on error This is what's done in oboe examples to avoid spurious errors breaking audio entirely.	2022-08-31 21:26:14 +01:00
Billy Laws	8917ec9c88	Don't set framesPerCallback for main stream as per oboe guidance It's best to let oboe figure it out on it's own	2022-08-31 21:26:14 +01:00
Billy Laws	b00008daf5	Fix identifier release check in AudioTrack::Stop	2022-08-31 21:26:14 +01:00
Billy Laws	d9c8e62d1c	Don't warn on GetConfig IOCTL fails	2022-08-31 21:26:14 +01:00
Billy Laws	4aef24ba32	Implement NVGPU_GPU_IOCTL_GET_GPU_TIME in nvdrv	2022-08-31 21:26:14 +01:00
Billy Laws	5841799420	Fix decoding of IOCTLs with padding at the end	2022-08-31 21:26:14 +01:00
Billy Laws	82444f3b0a	Don't set push descrptor flag for desc sets This is redundant and against the spec since we no longer use push descriptors.	2022-08-31 21:26:14 +01:00
PixelyIon	94fdd6aa43	Fix HID touch points not being removed from screen Tapping anything in titles that supported touch (such as Puyo Puyo Tetris or Sonic Mania) wouldn't work due to the first touch point never being removed from the screen, it is supposed to be removed after a 3 frame delay from the touch ending. This commit introduces a mechanism to "time-out" touch points which counts down during the shared memory updates and removes them from the screen after a specified timeout duration.	2022-08-31 23:43:02 +05:30
PixelyIon	70ad4498a2	Write HID LIFO entries at fixed intervals Certain titles depend on HID LIFO entries being written out at a fixed frequency rather than on actual state change, not doing this can lead to applications freezing till the LIFO is filled up to maximum size, this behavior is seen in Super Mario Odyssey. In other cases such as Metroid Dread, the game can run into race conditions that would lead to crashes, these were worked around by smashing a button during loading prior. This commit introduces a thread which sleeps and wakes up occasionally to write LIFO entries into HID shared memory at the desired frequencies. This alleviates any issues as it fills up the LIFO instantly and correctly emulates HID Shared Memory behavior expected by the guest. Co-authored-by: Narr the Reg <juangerman-13@hotmail.com>	2022-08-31 22:49:36 +05:30
PixelyIon	7966bfa9f6	Fix PI update `KThread::waiterMutex` deadlock It was determined that deadlocks inside `KThread::UpdatePriorityInheritance` would not only arise from the first level of locking with `waitingOn->waiterMutex` but also the second level of locking with `nextThread->waiterMutex` which has now also been fixed to fallback when facing contention.	2022-08-28 20:15:08 +05:30
Abandoned Cart	86f6fc510e	Remove printing result message from author There is no purpose in printing the same result twice, so any duplicate messages should instead allow the field to be truncated.	2022-08-27 18:54:27 +05:30
Abandoned Cart	1013857fc4	Refactor subtitle as author to remove subtitle Subtitle is no longer used, so instances have been rerouted to author. DataItem was also updated to reflect the removal of a subtitle.	2022-08-27 18:54:27 +05:30
Abandoned Cart	04cae942ea	Follow typical per-file detail formatting Format the details in the expected format of individual files (instead of a complete game) and move the code to match the updated placement.	2022-08-27 18:54:27 +05:30
lynxnb	5d6eaee301	Correctly save/restore ROM version to/from game entry cache PR #1758 introduced a bug where the game list would be entirely loaded every time the app was opened. This commit addresses that issue, which was caused by the `version` member of the cached game list being serialized to file (although incorrectly) but never actually read back when deserializing.	2022-08-21 20:24:36 +02:00
lynxnb	cfa5f0e030	Fix OSC alpha not changing on button press	2022-08-20 17:00:40 +02:00
KikiManjaro	8fb4e62c28	Add version information about rom Review: Co-authored-by: Niccolò Betto <niccolo.betto@gmail.com>	2022-08-20 13:48:07 +02:00
KikiManjaro	3407f6d530	Add OSC opacity adjustment	2022-08-20 13:46:17 +02:00
lynxnb	d129fb09cd	Android Manifest cleanup * Remove `package` from manifest and from activity prefixes, gradle `namespace` will be used instead * Removed deprecated `android.support.PARENT_ACTIVITY` metadata entries * Make `MainActivity` and `SettingsActivity` launched in `singleTop` mode to avoid unnecessary activity restarts while navigating the app	2022-08-17 15:33:53 +02:00
lynxnb	d128856c7d	Update target SDK version to Android 13	2022-08-17 12:29:46 +02:00
lynxnb	39f398f76b	Update Kotlin (1.7.10), NDK (25.0.8775105), AGP (7.2.2) and Kotlin deps	2022-08-17 12:28:31 +02:00
lynxnb	e9618d9e2c	Use pragma pack directions for tightly packing structs containing `u128` Using `__attribute__((packed))` doesn't work in new NDKs when a struct contains 128-bit integer members, likely because of a ndk/compiler bug. We now enclose the requiring structs in `#pragma pack` directives to tightly pack them.	2022-08-17 12:22:11 +02:00
lynxnb	c4bf92a49f	Fix Kotlin compilation errors from incorrect overloading of null-safe types	2022-08-17 12:16:26 +02:00
Billy Laws	bf491f71f9	Simplify blit helper shader vertex order	2022-08-10 15:43:16 +01:00
Billy Laws	c32bec071c	Adjust blit src{X,Y} to account for centred sampling before calling into helper shader Since the blit engine itself samples from pixel corners and the helper shader from pixel centres teh src coordinates need to be adjusted to avoid the helper shader wrapping round on the final column.	2022-08-10 15:39:37 +01:00
Billy Laws	08f36aac33	Enable hades vertex position input workaround for Adreno Caused crashes in any games using geometry shaders as by default hades uses the position builtin directly.	2022-08-08 18:09:00 +01:00
Billy Laws	04e7b684d2	Enable vertexPipelineStoresAndAtomics, fragmentStoresAndAtomics and shaderStorageImageWriteWithoutFormat Vulkan features Used by Xenoblade Chronicles DE	2022-08-08 17:43:18 +01:00
Billy Laws	390558c802	Add partial support for legacy attribute conversion We previously missed the hades pass for attribute conversion leading to crashes when games would attempt to use such an attribute. The hades pass for this isn't a proper fix however as it modifies the IR directly and will break if any of the previous stages in the pipeline change. Enable it to allow for games using them to at least have a chance at working. In the long term the pass will be reworked on the hades side to avoid modifying the IR in a way that can't be undone.	2022-08-08 17:43:18 +01:00
Billy Laws	540437b547	Fixup index buffer view caching We forgot to set the view size, which would end up forcing a view to be recreated with every call	2022-08-08 17:43:18 +01:00
Billy Laws	c966cd3b26	Prevent runtimeInfo vertex state from leaking into wrong shaders This vertex state must only be present for the last pipeline stage that touches vertices, if it is present for other stages it could result in incorrect behaviour like performing TFB in the fragment shader or flipping device coordinates twice.	2022-08-08 17:43:13 +01:00
Billy Laws	c52d3195cf	Ensure shader stage enable state matches pipeline stage enable state As the code was before, if we had a shader that was disabled and enabled again after without being invalidated the pipeline stage would stay disabled and break rendering.	2022-08-08 17:40:35 +01:00
Billy Laws	b1c669ba14	Always keep the VertexB shader stage enabled HW doesn't allow disabling the VertexB stage, enforce this in code.	2022-08-08 17:40:35 +01:00
lynxnb	d5174175d1	Implement indexed quads support We previously only supported non-indexed quads. Support for this is implemented by converting the index buffer at record time and pushing the result into the megabuffer, which is then used as the index buffer in the final draw command.	2022-08-08 17:40:35 +01:00
lynxnb	e6741642ba	Split out megabuffer allocation from pushing data The `Allocate` method allocates the given amount of space in a megabuffer chunk, returning a descriptor of the allocated region. This is useful for situations where you want to write directly to the megabuffer, avoiding the need for an intermediary buffer.	2022-08-08 17:40:35 +01:00
Billy Laws	cdc6a4628a	Enable VK uint8 indices feature when supported	2022-08-08 17:40:35 +01:00
Billy Laws	dccc86ea97	Implement transform feedback with VK_EXT_transform_feedback Tested to work in Xenoblade Chronicles DE, the code handles both hades varying input and buffer setup.	2022-08-08 17:40:35 +01:00
Billy Laws	06053d3caf	Rewrite Fermi 2D engine to use the blit helper shader Entirely rewrites the engine and interconnect code to take advantage of the subpixel and OOB blit support offered by the blit helper shader. The interconnect code is also cleaned up significantly with the 'context' naming being dropped due to potential conflicts with the 'context' from context lock	2022-08-08 17:40:35 +01:00
Billy Laws	395f665a13	Implement a system for helper shaders together with a simple blit shader It is desirable for us to use a shader for blits to allow easily emulating out of bounds blits and blits between different swizzled colour formats. The helper shader infrastructure is designed to be generic so it can be reused by any other helper shaders that we may need in the future.	2022-08-08 17:40:35 +01:00
Billy Laws	1da1698f90	Disable unused Vulkan HPP setters and smart handles	2022-08-08 14:57:44 +01:00
Billy Laws	f4e58a9238	Remove redundant synchost creating a new buffer	2022-08-08 14:57:44 +01:00
Billy Laws	11a8feb037	Correct nvdrv DMA copy class ID Was wrongly copy and pasted.	2022-08-08 14:57:44 +01:00
Billy Laws	13e7b54c61	Ensure failed IOCTLs are logged as a warning log	2022-08-08 14:57:44 +01:00
Billy Laws	eeb86a4f8a	Calculate renderArea from min(attachments.dimensions...) Vulkan doesn't support a renderArea larger than that of the smallest attachment	2022-08-08 14:57:44 +01:00
Billy Laws	9ea658d0ed	Don't throw on unsupported TIC formats These sometimes spuriously occur in games during transitions, to avoid crashing during them just use the null texture if they occur and log an error log	2022-08-08 14:57:44 +01:00
Billy Laws	856818c8eb	Emulate the 'None' mipfilter by adjusting LOD Borrowed this technique from yuzu since Vulkan has no direct equivalent	2022-08-08 14:57:44 +01:00
Billy Laws	9d50b6d0f7	Avoid locking presentation mutex in GetTransformHint This caused slowdown in Pokemon as it was being called every frame	2022-08-08 14:57:44 +01:00
Billy Laws	460e6c9c84	Use raw pointers to hold constant buffer views The constant destruction and creation of `BufferView`s in cbuf-heavy games showed up as a large chunk of the profiler. Fix this by taking advantage of the fact that constant buffer `BufferView`s are never deleted and always kept around in the cache to just return a pointer to them in the cache.	2022-08-08 14:54:57 +01:00
Billy Laws	6b2e84712b	Avoid race in nvdrv debug prints Looking up the device name without locking it could race with map insertions or deletions, so lock it to avoid that	2022-08-08 13:24:23 +01:00
Billy Laws	683cd594ad	Use a linear allocator for most per-execution GPU allocations Currently we heavily thrash the heap each draw, with malloc/free taking up about 10% of GPFIFOs execution time. Using a linear allocator for the main offenders of buffer usage callbacks and index/vertex state helps to reduce this to about 4%	2022-08-08 13:24:21 +01:00
Billy Laws	70eec5a414	Store delegate attached state within the delegate itself Avoids a costly map lookup for every AttachBuffer call, this was a serious bottleneck in SMO	2022-08-08 13:23:26 +01:00
Billy Laws	0268e1d5a0	Force a submit before any i2m engine writes We need traps to be inplace so we dont end up overwriting a resource that's being actively used by the current context without setting it to dirty	2022-08-08 13:22:37 +01:00

... 8 9 10 11 12 ...

2108 Commits